DeepAQI: A Vision-Based EfficientNet Framework for Air Quality Index Prediction from Environmental Metadata | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article DeepAQI: A Vision-Based EfficientNet Framework for Air Quality Index Prediction from Environmental Metadata Yash Mishra, Kedarnath senapati This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9062054/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract This research presents a deep learning–based framework for predicting the Air Quality Index (AQI) using outdoor webcam images from U.S. National Parks. Traditional AQI measurement relies on ground‑based air sensors, which require continuous calibration and are often geographically sparse. Recent advances in computer vision have motivated the exploration of image-based AQI estimation, enabling scalable, low‑cost, and real‑time monitoring in remote regions. Leveraging the publicly available NPS_AQI_DB, which includes more than 120,000 webcam images paired with corresponding pollutant data (O₃, SO₂, RH, temperature), we investigate both image‑only and multi‑modal (image + tabular) models for AQI regression. Our primary architecture is EfficientNet‑B0, chosen for its strong performance–efficiency trade‑off. The model is trained on augmented 224×224 images using AdamW optimization, mixed‑precision training, and robust preprocessing to handle missing and corrupted data. To enhance interpretability, Grad‑CAM visualizations highlight regions influencing AQI prediction, often corresponding to sky visibility, haze thickness, and lighting conditions. Additionally, we evaluate a two‑tower fusion model combining CNN features with meteorological variables, demonstrating improved stability across pollution categories. Experimental results show that the best-performing model achieves MAE ≈ 11.2 and RMSE ≈ 15.3 on the test set, reflecting competitive performance given the inherent noise of visual air estimation. Comprehensive evaluation includes residual analysis, calibration curves, AQI‑bin MAE breakdown, per‑park performance, and temporal error trends. These findings confirm that vision-based AQI prediction is feasible and can supplement traditional monitoring networks, especially in visually accessible but sensor-limited environments. Future work will explore temporal modeling, domain adaptation, and deployment on low-power edge devices. Full Text Additional Declarations The authors declare no competing interests. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9062054","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":602502436,"identity":"67c24a17-b96d-441f-95a2-f47a2c6a839b","order_by":0,"name":"Yash Mishra","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA80lEQVRIiWNgGAWjYHACNoaPfyR4+OUPHwByJGSI0sI4s8FGTnIGWwJICw9RWph5G9KMDW7wGIB4hLUYHD/+7AHvjsOJDbd7Pr+6UWPBw8B++OgGvFrO5JgbSJ45nNg45+w265xjQIfxpKXdwKvlQA6bhAHb4cRmhtxtxkA2D9A7Zvi1nH/+TCIBqKWNIeeZcc4/YrTcSDCTONiWZswjkcP8OLeNCC2SN96YSTacsZGT4DlmxpzbJ8HDRsgvfOfTn0n/qZDgsT/e/Phzzrc6OX72w8fwalE4gGCzSYBJfMpBQL4BwWb+QEj1KBgFo2AUjEwAAKmlTQZsW6kyAAAAAElFTkSuQmCC","orcid":"https://orcid.org/0000-0002-5004-1797","institution":"Nitk Surathkal","correspondingAuthor":true,"prefix":"","firstName":"Yash","middleName":"","lastName":"Mishra","suffix":""},{"id":602502437,"identity":"6e722038-3e1d-48a7-9b92-72987837bddc","order_by":1,"name":"Kedarnath senapati","email":"","orcid":"","institution":"Nitk Surathkal","correspondingAuthor":false,"prefix":"","firstName":"Kedarnath","middleName":"","lastName":"senapati","suffix":""}],"badges":[],"createdAt":"2026-03-08 05:17:24","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-9062054/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9062054/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":104781493,"identity":"6048fe1c-f1dc-4472-a751-36b05d37780e","added_by":"auto","created_at":"2026-03-17 07:55:47","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":821540,"visible":true,"origin":"","legend":"","description":"","filename":"MACSR5.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9062054/v1_covered_89a1b56c-d375-448b-b596-450a87871128.pdf"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003eDeepAQI: A Vision-Based EfficientNet Framework for Air Quality\u003c/p\u003e\n\u003cp\u003eIndex Prediction from Environmental Metadata\u003c/p\u003e","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"Nitk Surathkal","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-9062054/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9062054/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThis research presents a deep learning–based framework for predicting the Air Quality Index (AQI) using outdoor webcam images from U.S. National Parks. Traditional AQI measurement relies on ground‑based air sensors, which require continuous calibration and are often geographically sparse. Recent advances in computer vision have motivated the exploration of image-based AQI estimation, enabling scalable, low‑cost, and real‑time monitoring in remote regions. Leveraging the publicly available NPS_AQI_DB, which includes more than 120,000 webcam images paired with corresponding pollutant data (O₃, SO₂, RH, temperature), we investigate both image‑only and multi‑modal (image + tabular) models for AQI regression.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eOur primary architecture is EfficientNet‑B0, chosen for its strong performance–efficiency trade‑off. The model is trained on augmented 224×224 images using AdamW optimization, mixed‑precision training, and robust preprocessing to handle missing and corrupted data. To enhance interpretability, Grad‑CAM visualizations highlight regions influencing AQI prediction, often corresponding to sky visibility, haze thickness, and lighting conditions. Additionally, we evaluate a two‑tower fusion model combining CNN features with meteorological variables, demonstrating improved stability across pollution categories.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eExperimental results show that the best-performing model achieves MAE ≈ 11.2 and RMSE ≈ 15.3 on the test set, reflecting competitive performance given the inherent noise of visual air estimation. Comprehensive evaluation includes residual analysis, calibration curves, AQI‑bin MAE breakdown, per‑park performance, and temporal error trends. These findings confirm that vision-based AQI prediction is feasible and can supplement traditional monitoring networks, especially in visually accessible but sensor-limited environments. Future work will explore temporal modeling, domain adaptation, and deployment on low-power edge devices.\u003c/p\u003e","manuscriptTitle":"DeepAQI: A Vision-Based EfficientNet Framework for Air Quality\nIndex Prediction from Environmental Metadata","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-12 10:04:31","doi":"10.21203/rs.3.rs-9062054/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"a00d7c71-a66d-4dc7-ad16-9b75e7488a2d","owner":[],"postedDate":"March 12th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-03-12T10:04:31+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-12 10:04:31","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9062054","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9062054","identity":"rs-9062054","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.