A Lightweight, High-Throughput Classifier for North American Insects Using EfficientNet: Elytra 1.0

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Large-scale biodiversity monitoring is often inhibited by taxonomic obstacles. While deep learning has demonstrated efficacy in species identification, the increasing reliance on large Vision Transformers (ViTs) creates computational barriers that restrict usage to cloud-based infrastructure. Recent foundation models, such as BioCLIP and the Insect-1M framework, require parameter counts exceeding 100M, rendering them unsuitable for edge deployment in field operations. This study presents Elytra 1.0, a computer vision model optimized for edge deployment and capable of classifying 3,127 common North American insect species. The dataset, comprising 2.6 million images, includes all insect species in North America with over 1,000 research-grade observations on iNaturalist. An EfficientNet-B0 architecture was trained using transfer learning from ImageNet with adaptive learning rate scheduling. The model achieved 91.27% Top-1 Accuracy and 97.6% Top-5 Accuracy on an internal test set (N=289,151 images). To rigorously evaluate generalization beyond photographer-specific patterns, an independent observer-excluded test set (N=5,780 images, 578 species) was constructed comprising images exclusively from photographers who contributed zero training data. A post-hoc spatiotemporal audit revealed this test set was heavily skewed toward the Neotropics (Mean Lat: 6.05° N) during the boreal winter (Dec 2025–early Feb 2026). Despite this significant biogeographic and phenological shift from the predominantly temperate training data, the model achieved 86.68% Top-1 Accuracy (95% CI: 85.8–87.5%). This confirms that Elytra 1.0 relies on robust morphological features rather than learning background environmental correlations, maintaining high performance even in novel ecological contexts.The resulting model file size is 30 MB with an inference speed exceeding 700 frames per second (FPS) on mobile hardware. These results indicate that optimized convolutional architectures can achieve competitive accuracy with server-grade transformers while remaining suitable for decentralized, offline monitoring applications.
Full text 2,209 characters · extracted from oa-doi-fallback · click to expand
Abstract Large-scale biodiversity monitoring is often inhibited by taxonomic obstacles. While deep learning has demonstrated efficacy in species identification, the increasing reliance on large Vision Transformers (ViTs) creates computational barriers that restrict usage to cloud-based infrastructure. Recent foundation models, such as BioCLIP and the Insect-1M framework, require parameter counts exceeding 100M, rendering them unsuitable for edge deployment in field operations. This study presents Elytra 1.0, a computer vision model optimized for edge deployment and capable of classifying 3,127 common North American insect species. The dataset, comprising 2.6 million images, includes all insect species in North America with over 1,000 research-grade observations on iNaturalist. An EfficientNet-B0 architecture was trained using transfer learning from ImageNet with adaptive learning rate scheduling. The model achieved 91.27% Top-1 Accuracy and 97.6% Top-5 Accuracy on an internal test set (N=289,151 images). To rigorously evaluate generalization beyond photographer-specific patterns, an independent observer-excluded test set (N=5,780 images, 578 species) was constructed comprising images exclusively from photographers who contributed zero training data. A post-hoc spatiotemporal audit revealed this test set was heavily skewed toward the Neotropics (Mean Lat: 6.05° N) during the boreal winter (Dec 2025–early Feb 2026). Despite this significant biogeographic and phenological shift from the predominantly temperate training data, the model achieved 86.68% Top-1 Accuracy (95% CI: 85.8–87.5%). This confirms that Elytra 1.0 relies on robust morphological features rather than learning background environmental correlations, maintaining high performance even in novel ecological contexts.The resulting model file size is 30 MB with an inference speed exceeding 700 frames per second (FPS) on mobile hardware. These results indicate that optimized convolutional architectures can achieve competitive accuracy with server-grade transformers while remaining suitable for decentralized, offline monitoring applications. Competing Interest Statement The authors have declared no competing interest.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-27T02:00:06.600101+00:00
License: CC-BY-4.0