Deep Learning Guided Video Compression for Machine Vision Tasks

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract In the video compression industry, video compression tailored to machine vision tasks has recently emerged as a critical area of focus. Given the unique characteristics of machine vision, the current practice of directly employing conventional codecs reveals inefficiency, which requires compressing unnecessary regions. In this paper, we propose a framework that more aptly encodes video regions distinguished by machine vision to enhance coding efficiency. For that, the proposed framework consists of deep learning-based adaptive switch networks that guide the efficient coding tool for video encoding. Through the experiments, it is demonstrated that the proposed framework has superiority over the latest standardization project, video coding for machine benchmark, which achieves a Bjontegaard delta (BD)-rate gain of 5.91% on average and reaches up to a 19.51% BD-rate gain.
Full text 12,416 characters · extracted from preprint-html · click to expand
Deep Learning Guided Video Compression for Machine Vision Tasks | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Deep Learning Guided Video Compression for Machine Vision Tasks Aro Kim, Seung-taek Woo, Minho Park, Dong-hwi Kim, Hanshin Lim, and 3 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4346457/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 20 Sep, 2024 Read the published version in EURASIP Journal on Image and Video Processing → Version 1 posted 4 You are reading this latest preprint version Abstract In the video compression industry, video compression tailored to machine vision tasks has recently emerged as a critical area of focus. Given the unique characteristics of machine vision, the current practice of directly employing conventional codecs reveals inefficiency, which requires compressing unnecessary regions. In this paper, we propose a framework that more aptly encodes video regions distinguished by machine vision to enhance coding efficiency. For that, the proposed framework consists of deep learning-based adaptive switch networks that guide the efficient coding tool for video encoding. Through the experiments, it is demonstrated that the proposed framework has superiority over the latest standardization project, video coding for machine benchmark, which achieves a Bjontegaard delta (BD)-rate gain of 5.91% on average and reaches up to a 19.51% BD-rate gain. Video compression Video coding for machines Deep learning Full Text Cite Share Download PDF Status: Published Journal Publication published 20 Sep, 2024 Read the published version in EURASIP Journal on Image and Video Processing → Version 1 posted Reviewers agreed at journal 16 May, 2024 Reviewers invited by journal 16 May, 2024 Editor assigned by journal 01 May, 2024 First submitted to journal 29 Apr, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4346457","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":303383548,"identity":"6c0c5530-df50-4b35-b457-2cd08a25936e","order_by":0,"name":"Aro Kim","email":"","orcid":"","institution":"Kyungpook National University","correspondingAuthor":false,"prefix":"","firstName":"Aro","middleName":"","lastName":"Kim","suffix":""},{"id":303383549,"identity":"0188035d-792c-4b71-9db6-c969e697a300","order_by":1,"name":"Seung-taek Woo","email":"","orcid":"","institution":"Kyungpook National University","correspondingAuthor":false,"prefix":"","firstName":"Seung-taek","middleName":"","lastName":"Woo","suffix":""},{"id":303383550,"identity":"6446fb16-06cc-4339-bfd6-399cbe35b092","order_by":2,"name":"Minho Park","email":"","orcid":"","institution":"Kyungpook National University","correspondingAuthor":false,"prefix":"","firstName":"Minho","middleName":"","lastName":"Park","suffix":""},{"id":303383551,"identity":"783a36fa-7e45-493f-bad9-d0819739beda","order_by":3,"name":"Dong-hwi Kim","email":"","orcid":"","institution":"Kyungpook National University","correspondingAuthor":false,"prefix":"","firstName":"Dong-hwi","middleName":"","lastName":"Kim","suffix":""},{"id":303383552,"identity":"2660cd52-bfc4-4a76-9407-653fa3b897e5","order_by":4,"name":"Hanshin Lim","email":"","orcid":"","institution":"Electronics and Telecommunications Research Institute","correspondingAuthor":false,"prefix":"","firstName":"Hanshin","middleName":"","lastName":"Lim","suffix":""},{"id":303383553,"identity":"6bcfd446-70ee-47a7-8cdb-45762e2685bd","order_by":5,"name":"Soon-heung Jung","email":"","orcid":"","institution":"Electronics and Telecommunications Research Institute","correspondingAuthor":false,"prefix":"","firstName":"Soon-heung","middleName":"","lastName":"Jung","suffix":""},{"id":303383554,"identity":"ca5e2ebb-911d-4d38-9e51-0b79c1a9c75a","order_by":6,"name":"Sangwoon Kwak","email":"","orcid":"","institution":"Electronics and Telecommunications Research Institute","correspondingAuthor":false,"prefix":"","firstName":"Sangwoon","middleName":"","lastName":"Kwak","suffix":""},{"id":303383555,"identity":"9b9b4d3f-7da6-4613-9aa7-46145e02f797","order_by":7,"name":"Sang-hyo Park","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAu0lEQVRIiWNgGAWjYNCCAzYGUFYC0VrSSNdymAQt8g28Bz9XnDlvbC6RwPjhB0NaPkEtBgf4kiXP3LhtZjkjgVmyhyHHsoGgFgYeA8mGD7dtDG4kMEgzMFQYENIBdBiP8c+GD+dAWph/E6WF4QCPmWTDjQNmQC1sQFtyCGsxOMyXZtlwJtnY4MzDNssegzQiHNbee/hmwzE7ww3Hkw/f+FGRTITDmHlgLMYGUGgQA3gIKxkFo2AUjIIRDgC5STjyLZzPpgAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0002-7282-7686","institution":"Kyungpook National University","correspondingAuthor":true,"prefix":"","firstName":"Sang-hyo","middleName":"","lastName":"Park","suffix":""}],"badges":[],"createdAt":"2024-04-30 05:41:10","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4346457/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4346457/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s13640-024-00649-w","type":"published","date":"2024-09-20T15:58:05+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":65437578,"identity":"c3ad0837-35d2-463a-a7bf-e9404fb3ba1e","added_by":"auto","created_at":"2024-09-27 12:15:53","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":10776559,"visible":true,"origin":"","legend":"","description":"","filename":"DLGuidedVCMTasksv2.3.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4346457/v1_covered_ef1f2a33-a158-48b8-bf94-7c5f403ba19f.pdf"}],"financialInterests":"","formattedTitle":"Deep Learning Guided Video Compression for Machine Vision Tasks","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":true,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"eurasip-journal-on-image-and-video-processing","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"jivp","sideBox":"Learn more about [EURASIP Journal on Image and Video Processing](http://jivp-eurasipjournals.springeropen.com)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/jivp/default.aspx","title":"EURASIP Journal on Image and Video Processing","twitterHandle":"@SpringerEng","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"BMC/SO AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Video compression, Video coding for machines, Deep learning","lastPublishedDoi":"10.21203/rs.3.rs-4346457/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4346457/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eIn the video compression industry, video compression tailored to machine vision tasks has recently emerged as a critical area of focus. Given the unique characteristics of machine vision, the current practice of directly employing conventional codecs reveals inefficiency, which requires compressing unnecessary regions. In this paper, we propose a framework that more aptly encodes video regions distinguished by machine vision to enhance coding efficiency. For that, the proposed framework consists of deep learning-based adaptive switch networks that guide the efficient coding tool for video encoding. Through the experiments, it is demonstrated that the proposed framework has superiority over the latest standardization project, video coding for machine benchmark, which achieves a Bjontegaard delta (BD)-rate gain of 5.91% on average and reaches up to a 19.51% BD-rate gain.\u003c/p\u003e","manuscriptTitle":"Deep Learning Guided Video Compression for Machine Vision Tasks","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-05-27 03:23:45","doi":"10.21203/rs.3.rs-4346457/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewerAgreed","content":"","date":"2024-05-16T18:33:46+00:00","index":0,"fulltext":""},{"type":"reviewersInvited","content":"","date":"2024-05-16T18:30:43+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2024-05-01T04:42:49+00:00","index":"","fulltext":""},{"type":"submitted","content":"EURASIP Journal on Image and Video Processing","date":"2024-04-30T01:39:59+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"eurasip-journal-on-image-and-video-processing","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"jivp","sideBox":"Learn more about [EURASIP Journal on Image and Video Processing](http://jivp-eurasipjournals.springeropen.com)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/jivp/default.aspx","title":"EURASIP Journal on Image and Video Processing","twitterHandle":"@SpringerEng","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"BMC/SO AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"e9be4515-cee8-4104-b9a2-a5fb18fdd5dc","owner":[],"postedDate":"May 27th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2024-09-27T10:46:06+00:00","versionOfRecord":{"articleIdentity":"rs-4346457","link":"https://doi.org/10.1186/s13640-024-00649-w","journal":{"identity":"eurasip-journal-on-image-and-video-processing","isVorOnly":false,"title":"EURASIP Journal on Image and Video Processing"},"publishedOn":"2024-09-20 15:58:05","publishedOnDateReadable":"September 20th, 2024"},"versionCreatedAt":"2024-05-27 03:23:45","video":"","vorDoi":"10.1186/s13640-024-00649-w","vorDoiUrl":"https://doi.org/10.1186/s13640-024-00649-w","workflowStages":[]},"version":"v1","identity":"rs-4346457","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4346457","identity":"rs-4346457","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00