Video-Based Cattle Behavior Detection for Digital Twin Development in Precision Dairy Systems

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-15

This study developed and evaluated a video-based framework using YOLOv11 and TimeSformer to detect cows and classify seven behaviors, achieving 85.0% accuracy for digital twin applications in precision dairy systems.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-15 · read from full text

This preprint studied a video-based computer vision pipeline to detect and track individual dairy cattle and classify seven behaviors in commercial barn conditions, producing identity-persistent, temporally annotated outputs for digital-twin use. Using 4,964 annotated clips expanded to 9,600 via targeted augmentation, the authors combined YOLOv11 for detection with ByteTrack for tracking and compared SlowFast versus TimeSformer for behavior recognition; TimeSformer reached 85.0% overall accuracy (macro-F1 0.84) with 22.6 fps on an RTX A100. The model’s attention visualizations focused on anatomically relevant regions, and the structured outputs included cow ID, start–end times, durations, and confidence for downstream modeling and 3D visualization, though the paper is explicitly a non-peer-reviewed preprint and therefore has not undergone journal review. This paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract Digital twins in dairy systems require reliable behavioral inputs. We develop a video‑based framework that detects and tracks individual cows and classifies seven behaviors under commercial barn conditions. From 4,964 annotated clips, expanded to 9,600 through targeted augmentation, we couple YOLOv11 detection with ByteTrack for identity persistence and evaluate SlowFast versus TimeSformer for behavior recognition. TimeSformer achieved 85.0% overall accuracy (macro‑F1 0.84) and real‑time throughput of 22.6 fps on RTX A100 hardware. Attention visualizations concentrated on anatomically relevant regions (head/muzzle for feeding and drinking; torso/limbs for postures), supporting biological interpretability. Structured outputs (cow ID, start-end times, durations, confidence) enable downstream use in nutritional modeling and 3D digital‑twin visualization. The pipeline delivers continuous, per‑animal activity streams suitable for individualized nutrition, predictive health, and automated management, providing a practical behavioral layer for scalable dairy digital twins.
Full text 205,042 characters · extracted from preprint-html · click to expand
Video-Based Cattle Behavior Detection for Digital Twin Development in Precision Dairy Systems | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Video-Based Cattle Behavior Detection for Digital Twin Development in Precision Dairy Systems Shreya Rao, Eduardo Garcia, Suresh Neethirajan This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8032374/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 11 You are reading this latest preprint version Abstract Digital twins in dairy systems require reliable behavioral inputs. We develop a video‑based framework that detects and tracks individual cows and classifies seven behaviors under commercial barn conditions. From 4,964 annotated clips, expanded to 9,600 through targeted augmentation, we couple YOLOv11 detection with ByteTrack for identity persistence and evaluate SlowFast versus TimeSformer for behavior recognition. TimeSformer achieved 85.0% overall accuracy (macro‑F1 0.84) and real‑time throughput of 22.6 fps on RTX A100 hardware. Attention visualizations concentrated on anatomically relevant regions (head/muzzle for feeding and drinking; torso/limbs for postures), supporting biological interpretability. Structured outputs (cow ID, start-end times, durations, confidence) enable downstream use in nutritional modeling and 3D digital‑twin visualization. The pipeline delivers continuous, per‑animal activity streams suitable for individualized nutrition, predictive health, and automated management, providing a practical behavioral layer for scalable dairy digital twins. Biological sciences/Biological techniques Biological sciences/Computational biology and bioinformatics Physical sciences/Engineering Physical sciences/Mathematics and computing Biological sciences/Zoology Digital twin Cattle behavior detection Deep learning Computer vision Precision livestock farming Dairy monitoring Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 1. Introduction Digital twin technology represents a paradigm shift in precision livestock farming, transforming how dairy operations monitor, analyze, and optimize animal management through the creation of virtual replicas of physical farm entities 1 , 2 . These advanced systems integrate real-time sensor data, behavioral patterns, physiological parameters, and environmental conditions to enable continuous monitoring, predictive analytics, and automated decision-making at scales and levels of precision previously unattainable 3 , 4 . In the dairy sector, digital twins combine behavioral monitoring, mechanistic physiological modeling, nutritional requirement assessment, and environmental impact analysis to generate individualized management strategies that improve both animal welfare and production efficiency 5 . Empirical evidence supports these benefits: digital twin deployment can enhance feed conversion efficiency by 15–20%, reduce veterinary interventions by 25–30%, and cut greenhouse gas emissions per unit of milk by as much as 25% 6,2 . These outcomes illustrate the potential of digital twins as transformative tools for addressing the economic, environmental, and welfare challenges facing modern dairy production. Robust and continuous behavioral monitoring thus serves as the foundation upon which these digital twins operate, providing dynamic, high-frequency data to model individual cow states and management outcomes. A central prerequisite for digital twin functionality is accurate and continuous behavioral monitoring. In dairy cattle, five behavioral categories like feeding, drinking, lying, standing, and ruminating serve as critical indicators of metabolic balance, reproductive status, disease onset, and welfare status 7 . Feeding behavior is closely tied to dry matter intake and energy balance, and its measurement is essential for nutritional modeling and ration optimization. Rumination patterns provide direct insights into digestive function and metabolic efficiency, with deviations often preceding metabolic disorders such as acidosis, ketosis, or displaced abomasum 5 . Lying behavior is a proxy for cow comfort and lameness risk; research has established that 10–14 hours of daily lying is optimal for both productivity and welfare outcomes 8 . Drinking behavior, while less frequent, is vital for assessing hydration, feed palatability, and heat stress, all of which impact milk yield and welfare 9 . Even standing, often overlooked as a passive activity, can indicate restlessness, discomfort, or estrus expression when measured systematically. Automated recognition of these behaviors is therefore indispensable for creating the continuous, structured data streams that feed into digital twin models 1 , 3 . While the biological relevance of these behavioral metrics is well established, the challenge lies in capturing them continuously and objectively in commercial farm environments. Conventional behavioral monitoring methods, however, present significant limitations that restrict their scalability and accuracy. Manual observation, though capable of capturing fine-grained detail, is labor-intensive, subjective, and economically unsustainable for commercial herds that often number in the hundreds. Even among trained observers, inter-observer variability can reach 15–25%, reducing the reliability of datasets essential for predictive modeling 7 , 10 . Wearable sensor systems, while offering continuous monitoring, face a series of practical and welfare-related challenges. Device loss rates of 5–15% create gaps in longitudinal tracking. Batteries require replacement every 3–6 months, imposing labor costs and necessitating animal handling that can induce stress. Calibration drift reduces measurement reliability, while sensor placement may alter natural behaviors, potentially biasing the very data intended for welfare and productivity optimization 5 , 6 . These shortcomings underscore the need for alternative approaches capable of generating high-quality, scalable behavioral data. Computer vision has emerged as the most promising solution to these limitations, offering a scalable, non-invasive method for monitoring livestock at both individual and group levels. Barn-based video systems enable continuous monitoring of multiple animals simultaneously without interfering with their natural behaviors, while providing richer postural and temporal information than wearable sensors. Video analytics can capture body orientation, head movements, and activity transitions that are critical for distinguishing behaviors such as feeding versus ruminating. Importantly, video-based systems scale cost-effectively, monitoring dozens of animals with a single camera, and integrate easily with existing farm surveillance infrastructure ( 11 , 2 ). Furthermore, standardized output formats from computer vision pipelines facilitate integration with farm management software, nutritional modeling tools, and digital twin frameworks. Recent advances in deep learning architectures have dramatically improved the accuracy and robustness of behavior recognition systems. State-of-the-art models have achieved 85–95% accuracy in multi-class livestock behavior classification tasks 3 , 10 , 12 . Transformer-based models, originally developed for natural language processing and now adapted for video understanding, show promise for long-duration behaviors. For example, transformer-based classifiers have reached over 90% accuracy in beef cattle behavior recognition, outperforming convolutional networks in capturing spatiotemporal dynamics 10 . Recent advances in deep learning have improved livestock behavior recognition, with state-of-the-art models reporting 85–95% accuracy 3 , 9 , 11 . Transformer-based models show particular promise for long-duration behaviors 9 . Yet most published systems remain focused on isolated recognition rather than operational digital-twin workflows: identity persistence, structured outputs, and real-time performance under barn variability are rarely addressed. Bridging this gap requires unified pipelines that connect video analytics to nutritional and management models via standardized, digital-twin-ready behavior streams. Several challenges continue to constrain the scalability and robustness of current livestock video-based monitoring systems. Benchmark datasets are typically small (less than 1,000 clips) and collected under controlled conditions, limiting generalization to commercial barns. ( 10,13 ). The largest publicly available dairy cattle dataset, CBVD-5, contains only 687 clips, insufficient for training deep learning models with robust generalization capability. Moreover, many studies report only offline evaluation results, with real-time performance rarely demonstrated. Yet, for digital twins to function as decision-support tools, latency must remain below 200 ms to provide continuous updates for nutrition, health, and management interventions ( 1,11 ). Current systems often lack structured outputs, providing only categorical classifications rather than temporally annotated behavioral profiles linked to individual animal IDs. Such structured data streams are critical for downstream integration with nutritional models (e.g., NRC equations), predictive health algorithms, and farm management platforms ( 4,5 ). Finally, while transformer-based architectures such as TimeSformer have shown strong performance in other domains, their systematic evaluation in livestock contexts remains limited ( 3,10 ). These gaps highlight the need for comprehensive evaluation frameworks that incorporate real-world datasets, real-time performance, and standardized data outputs tailored to digital twin requirements. Beyond technical performance, practical deployment factors must also be considered. Barn environments present variable lighting, occlusion, and crowding, all of which complicate computer vision performance. Night-time monitoring often produces reduced accuracy due to low-light conditions, while occlusion from feeding barriers or overlapping animals can obscure key anatomical features necessary for classification. Robust systems must therefore incorporate augmentation strategies and multi-angle camera setups to generalize effectively across such variability. Interpretability is another crucial requirement for industry adoption. Attention mechanisms that highlight anatomically meaningful regions, such as the head for feeding or the legs for lying, provide confidence to farmers, veterinarians, and regulators that the models’ predictions are biologically valid ( 7,9 ). Beyond model performance, scalability and cost-effectiveness remain decisive factors for commercial uptake. Systems that function in real time on widely available GPU hardware without excessive computational demands are more likely to achieve practical deployment in dairy barns. Collectively, these considerations make clear that video-based behavior detection is not only a technical challenge but also a linchpin for the broader adoption of digital twins in dairy farming. By enabling continuous, individualized monitoring, computer vision provides the behavioral data streams required to drive mechanistic nutritional modeling, predictive health analytics, and environmental impact optimization. Without robust and scalable behavior detection systems, the promise of digital twins for livestock remains unattainable. This study directly addresses these gaps through the development of a comprehensive video-based cattle behavior detection system explicitly designed for digital twin integration. We present a large-scale dataset collected under authentic barn conditions, incorporating natural environmental variability. An optimized real-time processing pipeline is implemented, combining YOLO detection and ByteTrack tracking for persistent individual cow identification with comparative evaluation of state-of-the-art SlowFast and TimeSformer architectures for behavior classification. Structured outputs, including temporally annotated behavior logs with cow IDs and event durations, are generated in formats compatible with nutritional modeling frameworks and Unity-based visualization systems. By systematically benchmarking performance across models and validating interpretability through spatiotemporal attention analysis, this work establishes a robust technical foundation for integrating behavior detection into dairy digital twins. The objectives of this study are therefore to: (1) construct a large-scale video dataset of core cattle behaviors under commercial barn conditions; (2) develop a real-time multi-animal tracking and classification pipeline; (3) conduct a systematic comparison of SlowFast and TimeSformer architectures; and (4) generate structured, biologically interpretable outputs tailored for digital twin applications in precision dairy systems. 2. Materials and Methods 2.1 Experimental Environment and Data Collection 2.1.1 Farm Facility Description The study was conducted at the Ruminant Animal Centre (RAC) of Dalhousie University’s Agricultural Campus (Truro, Nova Scotia, Canada), a research facility dedicated to advanced dairy production, nutrition, and management studies. Each Holstein cow is tethered in an individual stall with a lying space, a front feed manger and an adjacent in-stall water bucket. Locomotion is constrained to standing and lying within the stall and the cows access feed and water without leaving the stall. Environmental control within the barn is achieved through a combination of natural and mechanical ventilation systems. Adjustable sidewall curtains allow air exchange based on external weather conditions, while chimneys with exhaust fans and circulation fans maintain airflow and temperature uniformity. During warmer months, additional ventilation fans are directed toward the cows to alleviate heat stress. The barn follows a 19-hour light and 5-hour dark cycle, with lights turning off at 10:00 p.m. and on at 3:00 a.m. to align with the milking schedule and support circadian rhythm balance. The RAC herd currently consists of approximately 80 Holstein dairy cows, of which 40 are actively lactating. The animals represent a range of parity and lactation stages, enabling balanced behavioral observations across physiological conditions. Milking is performed twice daily, at 4:30 a.m. and 4:00 p.m., consistent with commercial dairy practices in Atlantic Canada. Cows receive a total mixed ration (TMR) composed of grass silage, corn silage, straw, and concentrate. The formulation is adjusted regularly based on protein and energy analyses of the forage components and the production stage of each group; non-lactating and dry cows receive a lower-energy TMR variant to maintain optimal body condition. All management and feeding practices comply with Canadian Council on Animal Care (CCAC) guidelines and standard Canadian dairy production protocols. The RAC maintains stringent biosecurity and welfare standards, ensuring that all animals receive routine veterinary oversight, comfortable housing, and nutritional management aligned with NRC (2001) recommendations. All animal procedures were approved by the Dalhousie University Animal Care and Use Committee (Protocol #2024-026, approval date 16‑05‑2024) and complied with CCAC guidelines. An ARRIVE Essential 10 checklist is included in the Supplementary Information. 2.1.2 Multi-Camera Surveillance System A high-definition closed-circuit surveillance system was installed to enable continuous behavioral monitoring of the dairy herd. The system consisted of a total of seven Panasonic IP cameras, each configured to record at 1920 × 1080 resolution with a frame rate of 25–30 fps. Six of the units were Panasonic WV-S35302-F2L 2MP Outdoor Vandal Dome Cameras, equipped with 2.4 mm fixed lenses, infrared (IR) illumination for low-light monitoring, and integrated microphones for capturing ambient sound. These cameras are IP66 and IK10 rated for environmental durability and impact resistance, and are compliant with FIPS 140-2 Level 3 standards, ensuring secure data handling. One additional Panasonic WV-X15700-V2L 4K Outdoor Bullet Camera was installed in a high-activity zone. This unit featured a 4.3–8.6 mm motorized zoom lens and an embedded AI engine capable of supporting up to nine analytic applications simultaneously, enabling high-resolution tracking of fine-grained interactions. Cameras were mounted at heights of 3–4 meters using Panasonic WV-QWL500-W wall brackets and connected via Proterial 61337-8 CAT6A armored plenum-rated cables. Strategic placement at overhead and angled perspectives ensured coverage optimization, overlapping fields of view, and minimization of blind spots, particularly around feed bunks, water troughs, and lying stalls. The system provided continuous 24/7 monitoring across both day and night cycles, with infrared capabilities enabling uninterrupted observation. To safeguard against data loss, cameras were supported by an UltraTech 1000VA/600W uninterruptible power supply (UPS), ensuring reliability during power fluctuations. 2.2 Dataset Construction and Annotation Framework 2.2.1 Video Preprocessing Pipeline The raw surveillance footage obtained from the multi-camera system was initially subjected to a systematic preprocessing pipeline to ensure its suitability for downstream behavioral analysis. The video streams were first screened manually to identify behaviour rich segments containing clear examples of feeding, drinking, lying and standing. Segments with excessive occlusion, poor lighting or limited behavioural activity were excluded to maintain spatial and temporal consistency across the dataset. This filtering step followed established practices in large-scale livestock video analysis, where minimizing noise is essential for reliable behavioral inference 14 . This pipeline is depicted in Fig. 1 . Following quality filtering, long video sequences were segmented into shorter, behavior-specific clips using a semi-automated workflow. Cows were localized in each frame using a YOLOv11 based object detection model 15 , which was trained on a custom dataset of 2308 images, each labeled with bounding boxes around individual Holstein cows. The dataset was divided into training (1923 images; 83%), validation (193 images; 8%), and test (192 images; 8%) splits. Preprocessing included auto-orientation and extensive data augmentation was applied during training to improve robustness. Augmentation operations 16 included horizontal flips, random resized crops (0–12% zoom), small rotations between − 10° and + 10°, brightness adjustments (-15% to + 15%), contrast and exposure variations (-10% to + 10%), saturation adjustments (-25% to + 25%), Gaussian blur (up to 2.5 pixels), and additive noise applied to 0.1% of image pixels. Each training image generated three augmented variants, effectively expanding dataset diversity. The YOLOv11 model was trained for 90 epochs with a batch size of 16 and a learning rate of 0.01. Training was conducted in Google Colab Pro. The final model achieved high detection performance with mAP@50 = 0.994, precision = 0.982 and recall = 0.992, indicating reliable cow detection across varying barn conditions. To maintain the identity of individual cows across frames and throughout video segments, the detections were linked via the ByteTrack multi-object tracking algorithm 17 , which has been shown to outperform traditional identity-preserving trackers in complex agricultural scenes. This ensured that each cow maintained a consistent ID across frames, even during occlusion or group interactions. Once tracking was established, per-cow bounding boxes were cropped from each frame and resized to 224×224 pixels, producing standardized video clips corresponding to individual animals. All extracted video segments were standardized to a fixed duration of 10 seconds. This length was selected as a compromise between capturing short, transient actions such as drinking, and longer-duration behaviors such as lying or ruminating, which often span several minutes 10 . A fixed temporal window allowed for uniform sampling across the dataset and simplified downstream training procedures. The approach is consistent with recent work in animal behavior recognition, which emphasizes the importance of balancing clip length with the temporal resolution required to distinguish between different behavioral states 18 . As a result, the preprocessing pipeline produced a curated set of behavior-focused clips that maintained both ethological relevance and computational tractability for annotation and model development. The full workflow, from raw video to processed clips, is illustrated in Fig. 2 . 2.2.2 Behavioral Annotation Protocol Behavioral annotation was performed on the curated 10-second clips to classify cow activity into seven categories: standing and feeding, lying and feeding, drinking, lying, lying and ruminating, standing and ruminating, and standing. These behaviors were selected because they represent the dominant components of the daily activity budget of dairy cattle and are directly linked to productivity, health, and welfare outcomes 19 . Each behavior was defined according to established ethological criteria and verified against visual distinguishability standards to ensure consistent labeling. Feeding was annotated when a cow’s head was directed toward the feed bunk, with visible engagement in feed intake 10 . Drinking was assigned when the muzzle was in contact with or directly above the water trough, typically accompanied by head movements consistent with water ingestion 20 . Lying was defined by a resting posture in which the torso was in contact with the stall surface, with limbs folded beneath or alongside the body 21 . Standing was annotated when the cow maintained an upright posture with all four hooves in contact with the ground but without locomotion 20 . Ruminating was identified primarily through cyclical jaw movements associated with cud chewing, which typically occurred during lying but was also observed in stationary standing postures. These operational definitions ensured that categories were mutually exclusive and visually distinguishable in the recorded footage. To ensure annotation reliability, a dual-annotator protocol was implemented. Two trained annotators independently labeled all clips, and inter-rater agreement was calculated using Cohen’s kappa coefficient 22 , which consistently exceeded 0.95 across the dataset, indicating excellent agreement. In cases of disagreement, annotators engaged in consensus discussions to resolve inconsistencies. This multilayered annotation protocol combined ethological rigor, human oversight, and veterinary expertise, ensuring both biological accuracy and reproducibility of the dataset. 2.2.3 Dataset Characteristics The final curated dataset comprised 4,964 behavior clips, each standardized to a fixed length of 10 seconds prior to augmentation. The dataset spanned multiple temporal dimensions of variation. Recordings were collected 24/7, thereby capturing fluctuations in behavior associated with environmental changes such as temperature, humidity, and ventilation dynamics. In addition, the dataset encompassed both diurnal cycles and nocturnal activity patterns, supported by infrared camera functionality that enabled continuous monitoring during low-light conditions at night as well. Physiological variation was also represented, as cows at different lactation stages were included, thereby accounting for differences in activity budgets across productive and non-productive phases of the dairy cycle ( 23,24 ). Analysis of class distributions revealed a naturally imbalanced dataset, reflecting the time-allocation patterns typical of Holstein dairy cattle. Lying was the most frequent behavior, comprising 24.2% of the dataset, followed by standing (18.3%), feeding (15.4%), ruminating (12.1%), and drinking (3.8%). Such distributions align with established ethological research demonstrating that dairy cows spend a substantial proportion of their daily cycle resting or lying, while drinking occupies only a small fraction of total activity 25 . Although the raw dataset exhibited a natural imbalance across behavioral categories, such skewed distributions present a methodological challenge for machine learning, as minority classes like drinking and ruminating may be underrepresented in model training. To address this issue, the imbalance was explicitly corrected in subsequent preprocessing steps through targeted data augmentation ( 26,27 ). These augmentation strategies expanded the representation of rare behaviors while maintaining the integrity of majority classes, resulting in a more balanced dataset that better supports model generalization. By correcting imbalance at the data preparation stage, the training corpus was aligned both with the ecological diversity of cow behaviors and the computational requirements for robust classification. 2.3 Data Augmentation Strategy To address the limitations of the naturally imbalanced dataset and to improve the robustness of behavior recognition models, a structured data augmentation strategy was employed. Augmentation has been shown to enhance the generalization of deep learning models by synthetically expanding training data diversity while preserving ethological validity 16 . In this study, augmentation was applied across spatial, photometric, and temporal dimensions, followed by selective class balancing to mitigate underrepresentation of infrequent behaviors. 2.3.1 Augmentation Techniques Spatial augmentations were designed to reduce overfitting to fixed barn layouts and camera viewpoints. Random resized cropping was applied with scaling factors between 0.8 and 1.2, allowing the network to learn from slightly zoomed-in and zoomed-out perspectives. Horizontal flipping was incorporated to mimic mirrored viewpoints, while safe rotations limited to ± 15° introduced natural variability in orientation without compromising behavioral interpretability. These operations helped the model generalize across subtle positional and angular differences that arise from camera placement or cow movement. Photometric augmentations were introduced to address variability in lighting conditions and sensor noise. Brightness adjustments of ± 20% and contrast modifications of ± 15% simulated the natural fluctuations in illumination across day-night cycles and seasonal changes. Additionally, Gaussian noise with σ = 0.02 was added to replicate the visual distortions that occur under low-light or high-contrast recording conditions. By incorporating these variations, the model was trained to remain invariant to non-behavioral visual artifacts while focusing on essential cues for classification. Examples of spatial and photometric augmentations applied to cow images are shown in Fig. 3 . These include random cropping, rotations, brightness adjustments, and noise addition, which increase dataset diversity while preserving ethological interpretability. Temporal augmentations accounted for the inherent variability in behavioral tempo and duration. Frame sampling rates were varied to simulate different effective frame rates, thereby ensuring that the model learned representations that were robust to temporal resolution changes. In addition, clip duration modifications were performed within a safe margin around the standard 10-second window, allowing the model to handle both shorter clips (e.g., drinking events) and slightly extended sequences (e.g., lying or ruminating). These operations encouraged the network to capture temporal dynamics of behavior at multiple granularities, a practice consistent with best practices in spatio-temporal modeling 28 , 29 . 2.3.2 Class Balancing Implementation In addition to enhancing data diversity, augmentation was selectively employed to address the class imbalance observed in the raw dataset. Rare behaviors such as drinking and ruminating were deliberately oversampled through targeted augmentation factors, while more frequent behaviors were augmented conservatively. Specifically, drinking clips were augmented 7.5-fold, ruminating clips 4.5-fold, and the remaining behaviors (feeding, standing, lying) by 2-3-fold. This strategy expanded the dataset to over 9,600 clips, resulting in a substantially more balanced distribution across the behaviors. The final dataset composition included feeding (22.9%), standing (21.9%), lying (20.8%), ruminating (18.8%), and drinking (15.6%). By mitigating imbalance while preserving the ecological realism of behavior patterns, the augmented dataset provided a robust foundation for training behavior recognition models capable of performing reliably in real-world farm environments. The effect of the augmentation pipeline on class balance is illustrated in Fig. 4a and Fig. 4b, which compares the behavioral class distributions of the unaugmented and augmented datasets. As shown, augmentation substantially increased the representation of minority behaviors such as drinking and ruminating thereby improving dataset uniformity for model training. 2.4 Deep Learning Architecture Implementation To evaluate spatiotemporal modeling approaches for cow behavior recognition, two state-of-the-art video classification architectures were implemented: the SlowFast network 29 and the TimeSformer model 30 . These architectures were selected because they represent complementary strategies for capturing both short-term motion dynamics and long-range temporal dependencies, which are essential for distinguishing between rapid actions such as drinking and extended behaviors such as lying or ruminating. 2.4.1 SlowFast Network Configuration The SlowFast network employs a dual-pathway design, consisting of a slow pathway that processes frames at a low temporal resolution and a fast pathway that captures finer temporal granularity. In this study, the slow pathway operated on 8 frames per clip sampled at fixed intervals, while the fast pathway processed 32 frames per clip at a reduced spatial resolution of 224×224 pixels. This configuration allowed the slow pathway to capture global semantic context such as posture (e.g., lying vs. standing), while the fast pathway focused on short-term dynamics such as head movements during feeding or drinking. The two pathways were connected via lateral fusion layers, enabling information flow between slow and fast branches. These connections ensured that fine-grained motion features extracted by the fast branch enriched the semantic representations of the slow branch. Both pathways used a 3D convolutional backbone (ResNet-style), with spatiotemporal kernels applied across the frame sequence to jointly model motion and appearance. Computational efficiency was achieved by applying fewer channels in the fast pathway (1/8 of the slow pathway), reducing redundancy while preserving motion sensitivity. This design is particularly suited to livestock behavior analysis, as it mirrors the temporal scales of dairy cow activity budgets: fast dynamics (head/muzzle motion during feeding or drinking) and slow dynamics (postural changes such as lying or standing) 29 . 2.4.2 TimeSformer Architecture Design The TimeSformer model adopts a transformer-based architecture that replaces 3D convolutions with divided spatiotemporal self-attention mechanisms. Each input video clip was decomposed into non-overlapping patches, which were linearly embedded and fed into transformer encoder layers. Attention was computed separately along spatial and temporal dimensions, reducing computational cost compared to full joint attention while retaining modeling power 30 . In the spatial attention module, multi-head self-attention with 8 heads was applied to capture dependencies among different regions within each frame. In parallel, the temporal attention module modeled relationships across frames, allowing the network to capture long-range dependencies such as the transition from standing to lying or sustained ruminating bouts. To ensure position awareness, learned positional encodings were incorporated, enabling the model to distinguish between patches based on both spatial location and temporal order. A hybrid attention scheme was employed, in which layers alternated between spatial and temporal attention. This design provided the model with a global receptive field across both space and time, while maintaining computational tractability. By leveraging self-attention, the TimeSformer was able to integrate contextual information across the entire clip, making it especially effective for recognizing behaviors characterized by subtle, distributed cues. 2.5 Training Configuration and Optimization 2.5.1 Hardware and Software Setup All experiments were conducted on a cloud based GPU platform (Google LLC, USA), which provided access to an NVIDIA L4 GPU (24 GB VRAM) for model training. This cloud-based setup ensured sufficient memory and computational capacity for handling spatiotemporal video architectures such as 3D CNNs and transformers. Local preprocessing, annotation handling, and lightweight experiments were performed on a MacBook Air with Apple M3 chip (8-core CPU, integrated GPU, 16 GB unified memory). The computational environment was configured on Ubuntu 20.04 LTS with CUDA 11.6 and PyTorch 1.12.0 serving as the primary deep learning framework. The environment is also equipped with standard libraries for computer vision and video processing, including OpenCV 4.7, NumPy, and scikit-learn. This setup provided a reproducible and flexible platform for both development and large-scale training. 2.5.2 Model-Specific Training Parameters Model training was tailored to the architectural and computational requirements of each network. For the SlowFast network 29 , training was performed with a batch size of 4, using the Adam optimizer with a learning rate of 1 × 10⁻⁴. The loss function was cross-entropy, and the model was trained for 10 epochs, with extensions to 20 when necessary. Data loading employed two worker threads (num_workers = 2), with shuffling enabled to maximize diversity. For the TimeSformer model 30 , a more advanced optimization scheme was required to stabilize transformer training. Each input consisted of 12 frames of 224×224 resolution, processed in batches of 4 samples, with gradient accumulation across 4 steps to achieve an effective batch size of 16. Training proceeded for 20 epochs using the AdamW optimizer, with a learning rate of 3 × 10⁻⁵, weight decay of 0.01, and a warmup ratio of 0.1 to gradually ramp learning in the early iterations. To ensure numerical stability, gradients were clipped at a norm of 1.0, and mixed precision training was enabled to improve efficiency. Early stopping was applied with a patience of 4 epochs, halting training when validation performance failed to improve. 2.6 Real-Time Processing Pipeline 2.6.1 End-to-End System Architecture The proposed framework was implemented as a modular end-to-end video analysis pipeline, enabling automated behavior recognition from continuous surveillance footage of dairy cows. The system integrated three major components: object detection and tracking, behavior classification, and temporal smoothing. Raw video input from the multi-camera surveillance system was first processed through the YOLOv11 object detection model, trained on manually annotated cow images as described in Section 4.2.1. This stage localized individual cows within each frame. The resulting detections were then passed to the ByteTrack multi-object tracking algorithm, which maintained consistent cow identities across successive frames and prevented errors due to occlusions or group interactions. This ensured that each animal’s behavior was modeled continuously over time rather than as isolated instances. The cropped, per-cow video clips were subsequently fed into the behavior classification models (SlowFast or TimeSformer), which predicted one of seven behavioral states: standing and feeding, drinking, lying and feeding, lying, standing, standing and ruminating or lying and ruminating. To accommodate model input requirements, a sliding window strategy was employed, where clips of 16–32 consecutive frames were extracted with a 50% overlap and an 8-frame stride. This design captured sufficient temporal context for behavior recognition while ensuring efficient use of computational resources. Finally, to improve robustness against frame-level misclassifications, the system incorporated majority voting and temporal consistency enforcement. Predicted labels within overlapping windows were aggregated, and transitions were only accepted if they persisted across multiple consecutive windows. This reduced spurious label switching and produced smoother, ethologically plausible behavior timelines. By combining real-time detection, identity-preserving tracking, spatiotemporal classification, and temporal smoothing, the pipeline provided a reliable end-to-end solution for continuous monitoring of cow behavior in commercial farm environments. 2.6.2 Output Generation and Format Standards The final stage of the pipeline was responsible for converting model predictions into structured outputs suitable for downstream analysis, integration with nutritional modeling frameworks, and visualization in digital twin environments. To ensure both interoperability and reproducibility, standardized formats were adopted for data storage and streaming. All predictions were logged into structured CSV files containing the following fields: cow ID, predicted activity label, start timestamp, end timestamp, activity duration, and model confidence score. These logs enabled quantitative evaluation of behavioral activity budgets while preserving the temporal context of each prediction. Each entry was aligned with the synchronized video timeline, ensuring that results could be directly traced back to the original raw footage. To support domain-specific applications, the output schema was designed for compatibility with NRC-based nutritional modeling frameworks, where activity budgets (e.g., feeding or lying duration) inform energy expenditure and intake calculations. In addition, the logs were formatted for seamless integration with Unity-based visualization systems, where predicted states were used to drive cow avatars in a digital twin of the barn environment. This facilitated intuitive, real-time interpretation of behavioral patterns by both researchers and farm managers. For real-time applications, the system also incorporated low-latency streaming capabilities. Predictions were generated and transmitted in near real-time, with latency optimization achieved through GPU-accelerated inference, sliding window buffering, and asynchronous log writing. This ensured that activity updates were available within seconds of observation, enabling the framework to function not only as a retrospective analysis tool but also as a foundation for real-time monitoring and decision support systems in precision dairy farming. 3. Results 3.1 Dataset Characterization and Distribution Analysis 3.1.1 Pre- and Post-Augmentation Comparison The annotated dataset comprised seven distinct behavioral classes extracted from continuous barn recordings. As summarized in Fig. 4, the augmented dataset achieved a near-uniform class balance following the targeted oversampling strategy described in Section 4.3.2. Post-augmentation, the behavioral proportions aligned closely with known ethological activity budgets of Holstein cows lying and standing remained the most frequent states, while drinking and ruminating increased to biologically realistic levels. This balanced representation provided a stable foundation for downstream model training, ensuring that minority behaviors were adequately represented without distorting the natural behavioral hierarchy. 3.1.2 Natural Behavioral Frequency Analysis The overall behavioral distribution closely aligns with known diurnal activity patterns of Holstein cows under tie-stall housing conditions. Lying and standing behaviors were predominant during non-feeding hours, reflecting periods of rest and comfort 31 . Feeding activity peaked during scheduled feed deliveries (typically morning and late afternoon), while ruminating bouts were interspersed throughout the day, particularly following feeding episodes 23 . Drinking events occurred less frequently but were distributed evenly across the photoperiod 32 . These observations confirm that the collected dataset reflects biologically realistic behavioral proportions rather than sampling artifacts. 3.1.3 Augmentation Effectiveness Assessment The augmentation pipeline effectively expanded the training set size while preserving spatiotemporal consistency and behavioral realism. Augmented samples retained anatomical fidelity (e.g., head alignment with feed bunks or water troughs) and contextual cues (stall boundaries, lighting gradients) critical for accurate model training. Quantitatively, class balance improved from an initial max:min ratio of ~ 9:1 to approximately 2:1, reducing overfitting risk and improving minority-class recognition during preliminary validation runs. Collectively, these preprocessing and augmentation procedures produced a balanced, ecologically valid dataset capable of supporting robust training of deep learning models for fine-grained cow behavior recognition. 3.2 Model Performance Comparative Analysis The comparative evaluation of TimeSformer and SlowFast models aimed to determine the optimal spatiotemporal architecture for recognizing complex cattle behaviors under realistic barn conditions. Both models were trained and validated on the same curated seven-class dataset encompassing Drinking, Feeding and Lying, Feeding and Standing, Lying, Ruminating and Lying, Ruminating and Standing, and Standing. Each model was trained for 10 epochs using Adam optimization (learning rate = 1 × 10⁻⁴, batch size = 8) and identical augmentation pipelines (horizontal flip, temporal jittering, illumination normalization). The dataset was split into 70% training, 15% validation, and 15% testing sets. 3.2.1 Overall Classification Accuracy Across all evaluation runs, the TimeSformer achieved a mean overall accuracy of 85.0%, exceeding the SlowFast baseline accuracy of 82.3%. This improvement reflects the transformer’s inherent strength in capturing long-range temporal dependencies via its divided space-time self-attention mechanism. By encoding each video clip as a series of patch-level tokens, TimeSformer effectively attends to extended motion sequences such as continuous feeding or rumination cycles. The SlowFast model, although powerful in detecting high-frequency actions like drinking or head-lifting, relies primarily on convolutional hierarchies that emphasize localized motion and hence can miss broader behavioral context when frame-to-frame differences are minimal. To ensure the robustness of the observed performance gap, five independent trials were executed using different random initializations. The mean accuracy difference of 2.7 percentage points was statistically significant (p = 0.031, α = 0.05) based on a paired-sample t-test. This reproducibility underscores the stability of the transformer encoder for real-world deployment, where models must maintain reliable behavior detection despite day-to-day environmental variability. The superior accuracy of TimeSformer therefore validates the hypothesis that attention-driven global context modeling is essential for continuous barn surveillance, in which many cow behaviors manifest over extended time horizons rather than through abrupt motion cues alone. 3.2.2 Per-Class Performance Evaluation A fine-grained comparison of precision, recall, and F1-scores () provides further insight into class-specific performance trends. Table 1 summarizes the observed metrics and the corresponding improvements (ΔF1) obtained by the transformer-based model. Table 1 Per-class F1-scores for SlowFast and TimeSformer models across seven cattle behavior categories. The TimeSformer consistently outperformed the SlowFast network for most behaviors, particularly in Feeding & Standing, Standing, and Ruminating & Standing, reflecting its superior capacity to capture long-range temporal dependencies. Minor reductions for static postures (Lying and Ruminating & Lying) suggest that convolution-based architectures remain advantageous for low-motion states. Behavior Class SlowFast F1 TimeSformer F1 Drinking 0.910 0.901 Feeding & Lying 0.854 0.831 Feeding & Standing 0.924 0.947 Lying 0.888 0.878 Ruminating & Lying 0.806 0.786 Ruminating & Standing 0.742 0.786 Standing 0.667 0.759 Macro Average 0.827 0.861 Weighted Average 0.851 0.861 The largest improvement was recorded for Standing followed by Ruminating & Standing. These behaviors exhibit limited limb motion but prolonged duration, where the attention mechanism excels at integrating subtle temporal features - such as slight head tilts or jaw motions spread over many frames. Conversely, modest F1-score reductions for static postures (Lying, Ruminating & Lying) suggest that when motion cues are almost absent, TimeSformer’s patch-level embeddings may average out fine-grained micro-movements. Macro-averaged F1 improved from 0.827 (SlowFast) to 0.841 (TimeSformer), while the weighted F1 increased from 0.851 to 0.861. These gains confirm that transformer models generalize better across majority and minority behavior categories, yielding a more balanced classification under heterogeneous barn environments. Moreover, TimeSformer’s superior performance under nighttime and occluded conditions demonstrates its ability to learn illumination-invariant spatiotemporal representations, a crucial requirement for continuous, unattended surveillance in digital-twin pipelines. 3.3 Confusion Matrix and Error Analysis Whereas overall and per-class metrics quantify accuracy, confusion-matrix analysis exposes systematic misclassification patterns and the behavioral semantics behind them. To examine class-wise prediction behavior under different training conditions, confusion matrices were generated for both unaugmented and augmented datasets for each architecture. Figures 5(a-d) present a side-by-side comparison: (a) SlowFast - Unaugmented, (b) SlowFast - Augmented, (c) TimeSformer - Unaugmented, and (d) TimeSformer - Augmented. All matrices use identical color scaling for direct comparability. 3.3.1 Side-by-Side Confusion Matrix Comparison Both matrices exhibit strong diagonal dominance, confirming that most behaviors were correctly identified. However, persistent confusion clusters reveal key perceptual ambiguities inherent in cattle behavior recognition: Lying vs. Ruminating & Lying (~ 12–15% cross-confusion): These categories differ mainly through subtle jaw-movement cues. SlowFast’s motion-centric filters sometimes failed to detect small cyclic mandibular motions, labeling ruminating cows as merely lying. TimeSformer’s attention mechanism partially alleviated this overlap by tracking head-region temporal changes, thereby reducing off-diagonal errors. Standing vs. Ruminating & Standing: Both states share identical postural geometry, making visual separation challenging. TimeSformer achieved a higher Standing F1 (0.75 vs. 0.66) due to its capacity to integrate weak but consistent temporal signals, such as rhythmic chewing motions persisting across several seconds. Overall, the TimeSformer confusion matrix exhibits higher diagonal purity and reduced class-spillover, underscoring its superior ability to encode fine-grained temporal semantics. In practical terms, these improvements translate directly into more precise behavioral state transitions within the digital twin, ensuring that the simulated cow entities maintain realistic time-aligned activity cycles. 3.3.2 Error Source Attribution and Visual Inspection A post-hoc qualitative review of misclassified clips was performed to trace error origins and assess their implications for digital-twin fidelity. Three dominant sources were identified: Behavioral Ambiguity: Many errors stem from the continuum between similar activities. Static postures punctuated by micro-movements often blur class boundaries. For example, a cow transitioning from Lying to Ruminating & Lying may perform slow, imperceptible head adjustments that span multiple frames-challenging both convolutional and attention-based encoders. These findings emphasize the biological fluidity of behavior categories, suggesting that discrete classification may need to evolve toward probabilistic or continuous state modeling in future digital-twin iterations. Environmental and Lighting Variability: The barn environment introduced dynamic illumination, shadowing, and occlusion events. During dusk or under mixed natural-artificial lighting, contrast reduction led to occasional false positives in Drinking when reflective surfaces mimicked mouth-to-trough contact. Although TimeSformer exhibited stronger illumination resilience, both models showed mild accuracy drops in nocturnal segments, implying that domain-adaptive normalization or temporal brightness augmentation could further enhance generalization. Model Representation Bias: Minority classes (Feeding & Lying, Ruminating & Standing) contained fewer labeled instances, yielding under-optimized feature representations. The transformer’s tokenization may dilute critical motion vectors, while SlowFast’s motion-intensity bias sometimes over-fitted to fast activities. Mitigation strategies include focal loss weighting, synthetic clip generation, and attention-map regularization to improve the salience of under-represented actions. Implications for Digital-Twin Deployment: Misclassifications manifest as short-term state errors in the virtual twin, occasionally causing temporal discontinuities or overcounting of specific behaviors (e.g., fragmented rumination episodes). However, incorporating temporal smoothing, majority-vote state correction, and behavior-duration thresholds effectively attenuates these artifacts, preserving realistic simulation continuity. The qualitative review thus reinforces the importance of explainable model diagnostics-through attention-map visualization and per-frame attribution-to support transparent, trustworthy integration of computer-vision outputs into precision-nutrition digital twins. 3.3.3 Architecture-Specific Strengths and Weaknesses A comparative assessment of both models reveals complementary strengths arising from their distinct spatiotemporal encoding strategies. The TimeSformer, leveraging transformer-based self-attention, demonstrates superior sensitivity to feeding and drinking behaviors, where subtle head and mouth motions dominate. Its patch-token attention distributes weights dynamically across spatial regions, enabling the model to integrate micro-movements within longer temporal windows. Consequently, TimeSformer effectively identifies sequences where feeding occurs intermittently amid minimal body displacement, an ability crucial for capturing nutritional intake events in the digital twin. Conversely, the SlowFast architecture excels at posture driven behaviors such as Lying, Standing, and Ruminating & Lying. Its dual-pathway design where the fast branch encodes fine temporal granularity and the slow branch captures broader spatial context allows strong geometric posture recognition even under low motion conditions. The network’s convolutional inductive biases help stabilize predictions in cluttered environments, reducing false positives for large-scale static postures. Error-pattern inspection further reveals biologically consistent tendencies: misclassifications typically occur between physiologically adjacent states, such as transitions between Ruminating & Standing and Standing, or Lying and Ruminating & Lying. These correspond to authentic behavioral continuums rather than algorithmic artifacts, confirming that both models mirror the natural fluidity of bovine activities rather than imposing arbitrary separations. Hence, the observed confusions hold biological interpretability, underscoring the potential of deep spatiotemporal models as quantitative tools for ethological research. 3.4 Spatiotemporal Attention Visualization 3.4.1 Attention Heatmap Analysis To interpret model decision processes, attention-map visualizations 33 , 34 were generated for representative clips of each behavioral class. These visualizations reveal that learned focus areas align with anatomically relevant regions, reinforcing biological validity. During Feeding and Drinking, approximately 65% of cumulative attention weight concentrated around the head and muzzle region, particularly near the feed trough and water bucket interfaces. This indicates that the transformer successfully prioritized regions corresponding to ingestion activity, capturing the periodic forward-backward jaw motion characteristic of feeding bouts. For Standing and Lying postures, attention activation shifted toward body and leg contours, representing the model’s recognition of skeletal orientation and support distribution. In Ruminating & Lying states, attention maps alternated between head and thoracic regions, reflecting internalized understanding of chewing cycles within stationary frames. The corresponding attention visualizations are presented in Fig. 6 , which illustrate the spatial concentration of the model’s focus across representative behavioral classes. As shown, the TimeSformer consistently directs attention toward anatomically relevant regions such as the head and muzzle during feeding and drinking or the torso and limbs during postural states demonstrating biologically coherent feature learning. These spatial distributions confirm that model activations correspond to meaningful anatomical cues rather than spurious background patterns. By validating attention heatmaps against known ethological indicators-such as head movement frequency and body-weight distribution-the study establishes a direct link between learned visual salience and biological interpretability. This level of interpretability provides actionable transparency for veterinary decision support. By visualizing which regions and temporal segments drive classification, farm operators can corroborate algorithmic outputs with on-ground observations, improving trust in automated monitoring systems. In the context of the digital twin, attention transparency facilitates the translation of model activations into behavioral biomarkers, supporting nutritional adjustment decisions, welfare diagnostics, and anomaly detection. 3.5 Real-Time Processing Performance 3.5.1 Computational Efficiency Benchmarking Both architectures were benchmarked on an NVIDIA RTX A100 GPU on google colab using mixed-precision inference. The TimeSformer achieved an average throughput of 22.6 fps, surpassing SlowFast’s 18.4 fps due to its efficient token-wise computation and reduced temporal redundancy. Despite its transformer complexity, TimeSformer maintained lower GPU memory consumption (7.2 GB) compared to SlowFast’s 8.9 GB, attributed to the absence of multi-branch convolutional streams. From a deployment perspective, both models meet real-time processing thresholds for 25 fps video feeds 35 . However, the TimeSformer’s higher parallelization efficiency and consistent batch-to-batch latency (~ 44 ms/frame) make it preferable for continuous inference within digital-twin pipelines where computational scalability and stability are essential. A comparative hardware analysis indicates that real-time deployment is feasible on high-end consumer GPUs or edge-AI modules (e.g., Jetson AGX Orin) with slight down-sampling. Such resource profiling ensures that behavior recognition remains sustainable for on-farm automation without excessive power draw. 3.5.2 Scalability and Multi-Cow Processing To emulate barn-scale conditions, the end-to-end system was evaluated on simultaneous video streams tracking 16 individual cows using YOLOv11 + ByteTrack for detection and identity assignment. The combined detection-tracking-classification pipeline achieved an average latency of 180 ms per frame, corresponding to real-time throughput at 5.5 fps per cow. Batch-processing optimization and asynchronous GPU queues reduced total inference time by 27%, confirming the framework’s scalability to group monitoring without significant degradation in accuracy. These benchmarks demonstrate the system’s capacity for large-herd digital-twin synchronization, where multiple physical cows can be updated concurrently in the virtual environment. Future work will explore multi-GPU distribution and temporal batching to achieve near-real-time herd-level behavior analytics. 3.6 Challenge Analysis and Performance Limitations 3.6.1 Environmental Variability Impact Despite strong baseline accuracy, environmental factors continued to influence detection reliability. Under night-time illumination, accuracy declined by 8–12%, primarily due to reduced contrast and color-channel noise in infrared footage 9 . Occlusions-caused by barn structures, equipment, or cow overlapped to intermittent trajectory loss, occasionally truncating behavior sequences 3 . The tracking system successfully reidentified 92% of interrupted tracks, but brief identity swaps introduced minor temporal labeling noise. Corner and edge regions of the camera field exhibited degraded detection consistency, as cows partially exited the frame, limiting spatial continuity for both TimeSformer and SlowFast. Incorporating multi-view camera fusion and adaptive brightness equalization could mitigate these edge-case degradations. 3.6.2 Behavior-Specific Recognition Challenges Persistent confusion between Lying and Ruminating & Lying remains the principal classification bottleneck. Both behaviors share near-identical postural geometry; differentiation relies solely on subtle mandibular motion, which is occasionally obscured or temporally aliased at low frame rates 36 . Enhancing temporal resolution or integrating optical-flow-based motion cues could improve distinction in future iterations. Drinking behavior detection also posed challenges due to reflective water surfaces and variable head angles. False positives occurred when cows lowered their heads near troughs without actual ingestion, emphasizing the need for multi-modal fusion with acoustic or RFID-based water-intake sensors. Finally, temporal-sequence modeling limitations were evident in transitions between behaviors. Both architectures occasionally produced fragmented predictions during behavior shifts (e.g., standing to lying), suggesting that explicit temporal smoothing or recurrent attention integration could enhance continuity. These findings emphasize that while current deep-video architectures excel at single-state classification, the goal of continuous, biologically faithful behavioral tracking requires hybrid temporal models combining deep attention with probabilistic state transition logic. 4. Discussion The TimeSformer architecture demonstrated distinct advantages for livestock video analytics owing to its global spatiotemporal attention mechanism, which captures both spatial structures and long-term temporal evolution. By dynamically allocating attention to relevant regions such as the head, muzzle, or feed trough, the model maintains interpretability while adapting to motion sparsity typical of commercial barns. This allows robust recognition of prolonged behaviors like ruminating and feeding & standing, whose visual cues evolve slowly over time. In our 24/7 barn recordings, TimeSformer achieved 85.0% overall accuracy (macro-F1 = 0.84) and processed 22.6 fps on an RTX A100, confirming real-time feasibility. The SlowFast architecture, while slightly lower in accuracy (82.3%), remained competitive for posture-oriented states such as Standing or Lying. Its dual-pathway design (slow spatial, fast motion) efficiently captures short-duration transitions like drinking or standing changes, with predictable latency and low memory overhead, making it attractive for cost-constrained edge deployments. Compared with recent benchmarks CBVD-5 37 (687 segments, 107 Holsteins; ~ 78.7% accuracy) and the beef-cattle dataset of Cao et al. 10 (4974 clips; ~ 90% mAP₅₀) our dataset of 4964 annotated clips (expanded to 9600 via augmentation) encompasses seven behavior classes under natural lighting and occlusion, achieving parity with state-of-the-art accuracy despite far noisier conditions. The scale, behavioral granularity, and continuous capture across diurnal cycles make it one of the largest open dairy-behavior video resources, bridging the gap between controlled research corpora and operational barns. Nonetheless, several methodological limitations merit discussion. Identity leakage remains a primary risk if clips from the same cow appear in both training and test sets, apparent performance may be inflated. Future work will adopt leave-cow-out or day-blocked splits to ensure disjoint identities and report corresponding performance changes. Temporal sampling is another constraint 12 frames per 10-s clip (~ 1.2 fps) may undersample mandibular cycles critical for rumination and drinking, ablations at 16–32 frames will clarify the accuracy-latency trade-off. Tracking reliability also warrants quantification, our YOLOv11 and ByteTrack pipeline preserves per-cow IDs but has yet to be evaluated on standard multi-object tracking metrics (MOTA, IDF1, ID-switches) using an annotated subset. Environmental variability further affects performance; accuracy drops 8–12% at night or under severe occlusion, so stratified tables by day/night, view angle, and occlusion level will be added. Because augmented samples were confined to training, evaluation will be redone on unaugmented test data with natural class priors and balanced accuracy reporting to prevent optimistic bias. The behavioral taxonomy single-label composites such as “lying & feeding” or “standing & ruminating” simplifies ethograms that are inherently multi-label. These were chosen for practical monitoring in tie-stall barns, yet future work will re-analyze behavior along orthogonal posture and include confusion plots separating posture from oral-activity errors. The structured behavioral outputs cow ID, start/end times, durations, and confidence scores form the behavioral layer of the dairy digital twin. These CSV streams can directly feed NRC feed-intake models to infer dry-matter intake and dynamically adjust rations. Integration within Unity 3D enables real-time visualization of individual cows’ states synchronized with nutritional and environmental modules, establishing an operational link between perception and physiology. Within the twin dashboard, deviations from baseline feeding or rumination cycles act as early-warning indicators of metabolic or welfare issues, supporting proactive interventions. The architecture will be interoperable with commercial herd-management systems such as DairyComp 305 and Lely T4C and will be designed for modular deployment using networked edge devices. Both TimeSformer and SlowFast achieve real-time throughput on mid-range GPUs, ensuring cost-effective scalability. Economic analyses should consider reductions in manual observation labor, improved feed efficiency, and health-related cost savings. Reliability under barn humidity and dust will be maintained through ruggedized enclosures and modular camera design. Looking ahead, further improvements will emphasize multi-modal fusion, model calibration, and edge optimization. Acoustic and thermal sensors can complement visual data by capturing rumination chewing sounds and heat signatures, improving discrimination under occlusion or darkness. Lightweight transformer variants using pruning, quantization, and adaptive attention weighting will reduce inference cost and enhance interpretability. Federated-learning frameworks can enable cross-farm model updates without sharing raw video, preserving data privacy while addressing regional variability. Quantitative attention-map analyses reporting the proportion of attention mass within anatomical regions of interest will strengthen biological plausibility. Ultimately, integrating behavioral analytics with NRC feed-intake predictions, health metrics, and environmental sensing will create a self-updating digital-twin ecosystem that links perception, prediction, and management. By demonstrating continuous, per-cow behavioral timelines that synchronize physical animals with their virtual counterparts, this framework establishes a scalable foundation for individualized nutrition, early-warning health systems, and climate-smart, welfare-oriented dairy management. Declarations 5. Data Availability The annotated video dataset (4,964 clips) and augmentation recipes used to create the 9,600‑sample training set include representative sample clips are available on request. Due to facility restrictions, full raw CCTV footage is not publicly released; additional data are available from the corresponding author upon reasonable request. 6. Code Availability All scripts for preprocessing, YOLOv11 detection, ByteTrack tracking, and SlowFast/TimeSformer training and inference are available at https://github.com/mooanalytica/digital-twin-dairycow under the MIT license, with a frozen environment file and instructions to reproduce all reported results. 7. Author contributions S.R. curated data, implemented models, performed experiments, and wrote the manuscript. SR contributed to methodology and analysis. SN supervised the project, provided resources and critical revisions. All authors approved the final version. 8. Funding The authors sincerely thank the Natural Sciences and Engineering Research Council of Canada (RGPIN 2024-04450), the Net Zero Atlantic Canada Agency (300700018), Mitacs Canada (IT36514), and the Department of New Brunswick Agriculture, Aquaculture and Fisheries (NB2425-0025) for funding this study. 9. Competing interests The authors declare no competing financial or non‑financial interests. 10. Ethics statement All animal procedures were approved by the Dalhousie University Animal Care and Use Committee (Protocol #2024-026, approval date 16 May 2024) in accordance with Canadian Council on Animal Care (CCAC) guidelines. 11. Acknowledgments The authors gratefully acknowledge the staff and animal caretakers of the Ruminant Animal Centre at Dalhousie University for their assistance with data collection. References Youssef, A. et al. IUMENTA: A generic framework for animal digital twins within the Open Digital Twin Platform. Preprint at https://doi.org/10.48550/arXiv.2411.10466 (2024). Escribà-Gelonch, M. et al. Digital Twins in Agriculture: Orchestration and Applications. J Agric Food Chem 72 , 10737–10752 (2024). Zhang, Y. et al. Multimodal Behavior Recognition for Dairy Cow Digital Twin Construction Under Incomplete Modalities: A Modality Mapping Completion Network Approach. Preprint at https://doi.org/10.2139/ssrn.5097747 (2025). Zhang, Y. et al. Digital twin perception and modeling method for feeding behavior of dairy cows. Computers and Electronics in Agriculture 214 , 108181 (2023). Rao, S. & Neethirajan, S. Computational Architectures for Precision Dairy Nutrition Digital Twins: A Technical Review and Implementation Framework. Sensors 25 , 4899 (2025). Brown-Brandl, T. M. & Tao, J. ASAS-NANP Symposium: mathematical modeling in animal nutrition: harnessing real-time data and digital twins for precision livestock farming. Journal of Animal Science 103 , skaf138 (2025). Guarnido-Lopez, P., Pi, Y., Tao, J., Mendes, E. D. M. & Tedeschi, L. O. Computer vision algorithms to help decision-making in cattle production. Anim Fron 14 , 11–22 (2024). Kurras, F. & Jakob, M. Smart Dairy Farming—The Potential of the Automatic Monitoring of Dairy Cows’ Behaviour Using a 360-Degree Camera. Animals 14 , 640 (2024). Antognoli, V., Presutti, L., Bovo, M., Torreggiani, D. & Tassinari, P. Computer Vision in Dairy Farm Management: A Literature Review of Current Applications and Future Perspectives. Animals 15 , 2508 (2025). Cao, Z. et al. Semi-automated annotation for video-based beef cattle behavior recognition. Sci Rep 15 , (2025). Li, Z. et al. Method for Dairy Cow Target Detection and Tracking Based on Lightweight YOLO v11. Animals 15 , 2439 (2025). Bai, Q. et al. X3DFast model for classifying dairy cow behaviors based on a two-pathway architecture. Sci Rep 13 , 20519 (2023). Zia, A. et al. CVB: A Video Dataset of Cattle Visual Behaviors. Preprint at https://doi.org/10.48550/arXiv.2305.16555 (2023). Nasirahmadi, A., Edwards, S. A., Matheson, S. M. & Sturm, B. Using automated image analysis in pig behavioural research: Assessment of the influence of enrichment substrate provision on lying behaviour. Applied Animal Behaviour Science 196 , 30–35 (2017). Redmon, J., Divvala, S., Girshick, R. & Farhadi, A. You Only Look Once: Unified, Real-Time Object Detection. in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 779–788 (IEEE, Las Vegas, NV, USA, 2016). doi:10.1109/CVPR.2016.91. Shorten, C. & Khoshgoftaar, T. M. A survey on Image Data Augmentation for Deep Learning. J Big Data 6 , 60 (2019). Zhang, Y. et al. ByteTrack: Multi-object Tracking by Associating Every Detection Box. in Computer Vision – ECCV 2022 (eds Avidan, S., Brostow, G., Cissé, M., Farinella, G. M. & Hassner, T.) vol. 13682 1–21 (Springer Nature Switzerland, Cham, 2022). Wu, C., Zhou, Y., Pereia Pessôa, M. V., Peng, Q. & Tan, R. Conceptual digital twin modeling based on an integrated five-dimensional framework and TRIZ function model. Journal of Manufacturing Systems 58 , 79–93 (2021). Grant, R. J. & Albright, J. L. Effect of Animal Grouping on Feeding Behavior and Intake of Dairy Cattle. Journal of Dairy Science 84 , E156–E163 (2001). Yu, R. et al. Research on Automatic Recognition of Dairy Cow Daily Behaviors Based on Deep Learning. Animals (Basel) 14 , 458 (2024). Borchers, M. R., Chang, Y. M., Tsai, I. C., Wadsworth, B. A. & Bewley, J. M. A validation of technologies monitoring dairy cow feeding, ruminating, and lying behaviors. Journal of Dairy Science 99 , 7458–7466 (2016). McHugh, M. L. Interrater reliability: the kappa statistic. Biochem Med (Zagreb) 22 , 276–282 (2012). DeVries, T. J., Von Keyserlingk, M. A. G., Weary, D. M. & Beauchemin, K. A. Measuring the Feeding Behavior of Lactating Dairy Cows in Early to Peak Lactation. Journal of Dairy Science 86 , 3354–3361 (2003). Beauchemin, K. A. Invited review: Current perspectives on eating and rumination activity in dairy cows. Journal of Dairy Science 101 , 4762–4784 (2018). Tucker, C. B., Jensen, M. B., De Passillé, A. M., Hänninen, L. & Rushen, J. Invited review: Lying time and the welfare of dairy cows. Journal of Dairy Science 104 , 20–46 (2021). Jafarigol, E., Trafalis, T. & Mohammadi, N. A Review of Machine Learning Techniques in Imbalanced Data and Future Trends. Preprint at https://doi.org/10.48550/arXiv.2310.07917 (2025). Miftahushudur, T., Sahin, H. M., Grieve, B. & Yin, H. A Survey of Methods for Addressing Imbalance Data Problems in Agriculture Applications. Remote Sensing 17 , 454 (2025). Carreira, J. & Zisserman, A. Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. Preprint at https://doi.org/10.48550/arXiv.1705.07750 (2018). Feichtenhofer, C., Fan, H., Malik, J. & He, K. SlowFast Networks for Video Recognition. in 2019 IEEE/CVF International Conference on Computer Vision (ICCV) 6201–6210 (IEEE, Seoul, Korea (South), 2019). doi:10.1109/ICCV.2019.00630. Bertasius, G., Wang, H. & Torresani, L. Is Space-Time Attention All You Need for Video Understanding? Preprint at https://doi.org/10.48550/arXiv.2102.05095 (2021). Ito, K. ASSESSING COW COMFORT USING LYING BEHAVIOUR AND LAMENESS. Cardot, V., Le Roux, Y. & Jurjanz, S. Drinking Behavior of Lactating Dairy Cows and Prediction of Their Water Intake. Journal of Dairy Science 91 , 2257–2264 (2008). An, J. & Joe, I. Attention Map-Guided Visual Explanations for Deep Neural Networks. Applied Sciences 12 , 3846 (2022). Chefer, H., Gur, S. & Wolf, L. Transformer Interpretability Beyond Attention Visualization. in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 782–791 (IEEE, Nashville, TN, USA, 2021). doi:10.1109/CVPR46437.2021.00084. Li, Z. et al. Method for Dairy Cow Target Detection and Tracking Based on Lightweight YOLO v11. Animals 15 , 2439 (2025). Beauchemin, K. A. Invited review: Current perspectives on eating and rumination activity in dairy cows. Journal of Dairy Science 101 , 4762–4784 (2018). Li, K., Fan, D., Wu, H. & Zhao, A. CBVD-5 (Cow Behavior Video Dataset). Kaggle (2024). Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 01 Jan, 2026 Reviews received at journal 17 Dec, 2025 Reviewers agreed at journal 10 Dec, 2025 Reviews received at journal 24 Nov, 2025 Reviews received at journal 19 Nov, 2025 Reviewers agreed at journal 18 Nov, 2025 Reviewers agreed at journal 15 Nov, 2025 Reviewers invited by journal 13 Nov, 2025 Editor assigned by journal 12 Nov, 2025 Submission checks completed at journal 10 Nov, 2025 First submitted to journal 04 Nov, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8032374","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":547842641,"identity":"99a6b8fe-bddc-4bac-9d7f-0ba48d8f7ef6","order_by":0,"name":"Shreya Rao","email":"","orcid":"","institution":"Dalhousie University","correspondingAuthor":false,"prefix":"","firstName":"Shreya","middleName":"","lastName":"Rao","suffix":""},{"id":547842642,"identity":"86b11088-b65c-4b34-8417-37761d64d22b","order_by":1,"name":"Eduardo Garcia","email":"","orcid":"","institution":"Dalhousie University","correspondingAuthor":false,"prefix":"","firstName":"Eduardo","middleName":"","lastName":"Garcia","suffix":""},{"id":547842643,"identity":"8bd501a6-36e8-415b-b25d-ea35627180b4","order_by":2,"name":"Suresh Neethirajan","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA0klEQVRIiWNgGAWjYLACxgYQZj4gwVBAmha2BAkGA1K0MDDwGBCnRX5G7sOHP3fck2duP/Pxxg8DBnn+BgJaDG6kGxvznik2bOzJ3WzZY8BgOOMAIS0SaWzSjG0JjI0NudskeAwYEhgIaZGfkcb+82dbgn1j/5tnkn+AWuQJaWG4kcbGwNuWkNg4I4dNGmSLAUGHnXnGLA3Uktw445mxtYyBhOFGgg5rT2P8CHSY7cb+5Ic331TYyMsRdBgMGDaAKQli1YOsI0HtKBgFo2AUjDAAAEvwPwjEzMvaAAAAAElFTkSuQmCC","orcid":"","institution":"Dalhousie University","correspondingAuthor":true,"prefix":"","firstName":"Suresh","middleName":"","lastName":"Neethirajan","suffix":""}],"badges":[],"createdAt":"2025-11-04 21:08:12","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8032374/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8032374/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":96629507,"identity":"2f595be1-aa69-4425-8ada-21e369b219e5","added_by":"auto","created_at":"2025-11-24 12:28:11","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":8198665,"visible":true,"origin":"","legend":"","description":"","filename":"ShreyaPaperNov72025npj.docx","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/f80e5e715160b0383ac5398e.docx"},{"id":96708710,"identity":"4909a039-e74d-4597-adba-268a697fae2c","added_by":"auto","created_at":"2025-11-25 10:05:12","extension":"svg","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":634532,"visible":true,"origin":"","legend":"","description":"","filename":"Figure1.svg","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/664f2872e4cb0141cdba3bbe.svg"},{"id":96709352,"identity":"e5869cd2-bfc0-4440-948a-c0fd32ff272c","added_by":"auto","created_at":"2025-11-25 10:08:47","extension":"svg","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":5180810,"visible":true,"origin":"","legend":"","description":"","filename":"Figure2.svg","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/900a501b28454afe22e14528.svg"},{"id":96629494,"identity":"6f3675d2-8dea-4a46-9fc0-73696585de15","added_by":"auto","created_at":"2025-11-24 12:28:11","extension":"svg","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":174819,"visible":true,"origin":"","legend":"","description":"","filename":"Figure3.svg","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/b5576d40cb15d814fd29c97f.svg"},{"id":96709839,"identity":"ad319738-351b-420d-a57d-0f8d9ca904d3","added_by":"auto","created_at":"2025-11-25 10:09:43","extension":"svg","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":28832,"visible":true,"origin":"","legend":"","description":"","filename":"Figure4a.svg","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/0fc863b7c6e33e03ea1c5379.svg"},{"id":96629493,"identity":"63a72495-0b2d-422e-b479-613c58d58b24","added_by":"auto","created_at":"2025-11-24 12:28:11","extension":"svg","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":35184,"visible":true,"origin":"","legend":"","description":"","filename":"Figure4b.svg","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/f400298d51989ab48188b249.svg"},{"id":96709812,"identity":"5d62f5c1-a055-4982-bd6d-a53d675ae9b5","added_by":"auto","created_at":"2025-11-25 10:09:42","extension":"svg","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":89960,"visible":true,"origin":"","legend":"","description":"","filename":"Figure5a.svg","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/392b9c20212a291541a01ad4.svg"},{"id":96629506,"identity":"f90c14b9-8ccc-4c6b-a232-4b49ebbe5323","added_by":"auto","created_at":"2025-11-24 12:28:11","extension":"svg","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":86196,"visible":true,"origin":"","legend":"","description":"","filename":"Figure5b.svg","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/08302dd95fdb5739e6b419f5.svg"},{"id":96629503,"identity":"239a8545-0207-4be3-9759-f9f95695952e","added_by":"auto","created_at":"2025-11-24 12:28:11","extension":"svg","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":80136,"visible":true,"origin":"","legend":"","description":"","filename":"Figure5c.svg","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/837038ec888dc70ecc149610.svg"},{"id":96708859,"identity":"cbe93189-7206-40ba-8297-86db9e6a05cf","added_by":"auto","created_at":"2025-11-25 10:05:41","extension":"svg","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":25552,"visible":true,"origin":"","legend":"","description":"","filename":"Figure5d.svg","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/403ecd1c2b8bff2a2a45ce41.svg"},{"id":96709140,"identity":"477066b6-c0c2-4c6a-ac47-bc2e66667467","added_by":"auto","created_at":"2025-11-25 10:07:56","extension":"svg","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":10097,"visible":true,"origin":"","legend":"","description":"","filename":"Figure6a.svg","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/5d752edcec84248744633507.svg"},{"id":96710100,"identity":"04ccb846-7ea9-452f-9b44-1ed15fcc2e1e","added_by":"auto","created_at":"2025-11-25 10:10:04","extension":"svg","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":7957,"visible":true,"origin":"","legend":"","description":"","filename":"Figure6b.svg","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/c454ca577b330ac5aa640166.svg"},{"id":96708905,"identity":"9a94f75a-23d7-437e-b74f-d72b4ceeab69","added_by":"auto","created_at":"2025-11-25 10:06:08","extension":"json","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":4861,"visible":true,"origin":"","legend":"","description":"","filename":"3a8104b9211d4853998fb42ede7887a0.json","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/e438ee11d8c3ef683147bdd1.json"},{"id":96629498,"identity":"bde1f140-345d-442f-b1fc-344ee0bbf83c","added_by":"auto","created_at":"2025-11-24 12:28:11","extension":"xml","order_by":13,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":138953,"visible":true,"origin":"","legend":"","description":"","filename":"3a8104b9211d4853998fb42ede7887a01enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/b689f49cfbe92329eb61c0e5.xml"},{"id":96708980,"identity":"3e449b4c-909f-49ae-8d68-23e308081ff0","added_by":"auto","created_at":"2025-11-25 10:06:48","extension":"svg","order_by":14,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":634532,"visible":true,"origin":"","legend":"","description":"","filename":"Figure1.svg","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/bc4a7a673098962c556de73e.svg"},{"id":96629528,"identity":"5718355e-badc-4d27-bb00-69091f3b3fa7","added_by":"auto","created_at":"2025-11-24 12:28:12","extension":"svg","order_by":15,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":5180810,"visible":true,"origin":"","legend":"","description":"","filename":"Figure2.svg","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/9a0663ab6aa18d6f33ef2369.svg"},{"id":96629505,"identity":"ea144cc6-e05d-4423-8ac9-2f2d8cf1a482","added_by":"auto","created_at":"2025-11-24 12:28:11","extension":"svg","order_by":16,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":174819,"visible":true,"origin":"","legend":"","description":"","filename":"Figure3.svg","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/f8c339be08bdab8e077a2d49.svg"},{"id":96708761,"identity":"5b7f0497-f63b-49b2-8757-7e1c51838023","added_by":"auto","created_at":"2025-11-25 10:05:24","extension":"svg","order_by":17,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":28832,"visible":true,"origin":"","legend":"","description":"","filename":"Figure4a.svg","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/cfa3006e4361029bea428297.svg"},{"id":96629509,"identity":"3d557ddd-5d9e-483c-9dfb-c8f9b604c696","added_by":"auto","created_at":"2025-11-24 12:28:11","extension":"svg","order_by":18,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":35184,"visible":true,"origin":"","legend":"","description":"","filename":"Figure4b.svg","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/fcdd19c016f0b5f4d5160ee8.svg"},{"id":96709335,"identity":"71c86c23-4d87-4e6d-bbbf-6f7b76309532","added_by":"auto","created_at":"2025-11-25 10:08:45","extension":"svg","order_by":19,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":89960,"visible":true,"origin":"","legend":"","description":"","filename":"Figure5a.svg","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/ca012d2bc6bb43bb7d6e768f.svg"},{"id":96629532,"identity":"2c3c19a5-11db-48c0-95c5-0cd0462226a6","added_by":"auto","created_at":"2025-11-24 12:28:12","extension":"svg","order_by":20,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":86196,"visible":true,"origin":"","legend":"","description":"","filename":"Figure5b.svg","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/baa8408970d86b6ae75b2cf5.svg"},{"id":96709013,"identity":"eddc8eed-3b16-4808-be3f-22b462a341cc","added_by":"auto","created_at":"2025-11-25 10:07:06","extension":"svg","order_by":21,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":80136,"visible":true,"origin":"","legend":"","description":"","filename":"Figure5c.svg","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/3af83200a9335ae2e0e712d4.svg"},{"id":96629515,"identity":"e312ffdd-785c-4bc8-bc47-66779b6a2bb1","added_by":"auto","created_at":"2025-11-24 12:28:12","extension":"svg","order_by":22,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":25552,"visible":true,"origin":"","legend":"","description":"","filename":"Figure5d.svg","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/f76c6780a8ce8a971e465b5e.svg"},{"id":96629511,"identity":"9326f76b-2938-48d2-8060-217c4d5ae2c2","added_by":"auto","created_at":"2025-11-24 12:28:12","extension":"svg","order_by":23,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":10097,"visible":true,"origin":"","legend":"","description":"","filename":"Figure6a.svg","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/522b9f7cf940900b204f9f7e.svg"},{"id":96708781,"identity":"0fccbaa5-365b-437c-9100-62f0dc4f2634","added_by":"auto","created_at":"2025-11-25 10:05:26","extension":"svg","order_by":24,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":7957,"visible":true,"origin":"","legend":"","description":"","filename":"Figure6b.svg","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/7185a7de6d881aeed66ddc9f.svg"},{"id":96708551,"identity":"9f2450c3-de60-41e6-9f58-1f21ce4b981c","added_by":"auto","created_at":"2025-11-25 10:04:26","extension":"png","order_by":25,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":166781,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/0a23c8360e2aa5205d2db4ee.png"},{"id":96629522,"identity":"cbddc9bd-8aa0-4312-a631-3231b350aac6","added_by":"auto","created_at":"2025-11-24 12:28:12","extension":"jpeg","order_by":26,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":479676,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage10.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/8fd308e00241f2e0051db4dd.jpeg"},{"id":96709354,"identity":"2cc81e2a-bf6c-4e1f-abbc-d4d0c8903978","added_by":"auto","created_at":"2025-11-25 10:08:48","extension":"jpeg","order_by":27,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":363448,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/68ab483096a5bae7b8546a8e.jpeg"},{"id":96708955,"identity":"747d2e5e-4990-4ed1-8c44-cb6d00b98ba5","added_by":"auto","created_at":"2025-11-25 10:06:38","extension":"png","order_by":28,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":679616,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/b329af39d610ea37dfdca6fa.png"},{"id":96629520,"identity":"9ce2508f-c0ca-4324-83a9-c63d14771f88","added_by":"auto","created_at":"2025-11-24 12:28:12","extension":"png","order_by":29,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":21204,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/218ab5c205fe29ae227aafe8.png"},{"id":96709027,"identity":"437b1e74-9995-40ed-94bb-86cd4ec85865","added_by":"auto","created_at":"2025-11-25 10:07:12","extension":"png","order_by":30,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":23499,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/5f3d7518574afc59a0f6222f.png"},{"id":96709808,"identity":"464a56b3-f142-436c-927e-4e133f141c8c","added_by":"auto","created_at":"2025-11-25 10:09:42","extension":"png","order_by":31,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":93833,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/be7fa8bd6d22bb584de1ebed.png"},{"id":96629527,"identity":"412898a5-57a0-488f-846c-b0383ba9fa92","added_by":"auto","created_at":"2025-11-24 12:28:12","extension":"png","order_by":32,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":102110,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/7537757a199e125645c01c21.png"},{"id":96709319,"identity":"9226b0ef-21e5-481c-a5f0-95b601f47bf0","added_by":"auto","created_at":"2025-11-25 10:08:43","extension":"png","order_by":33,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":90787,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage8.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/b47471295286ee10561a8bf1.png"},{"id":96629521,"identity":"3b537f9c-0ac5-4c6e-bde1-7f0c7bdbaec2","added_by":"auto","created_at":"2025-11-24 12:28:12","extension":"png","order_by":34,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":58428,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage9.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/bcf211a307fb0e5a98203bfb.png"},{"id":96629534,"identity":"430bf065-5961-4cd9-975d-d49cad2cb297","added_by":"auto","created_at":"2025-11-24 12:28:12","extension":"png","order_by":35,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":100440,"visible":true,"origin":"","legend":"","description":"","filename":"OnlineFigure1.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/a3ffc0c3da94e930fd6b13b1.png"},{"id":96629524,"identity":"52b62336-d7f4-411e-9246-4f0a34271943","added_by":"auto","created_at":"2025-11-24 12:28:12","extension":"png","order_by":36,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1026476,"visible":true,"origin":"","legend":"","description":"","filename":"OnlineFigure2.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/6913f7c2fdcd228ab43ccf36.png"},{"id":96629525,"identity":"9a0f6e4d-346f-4012-94e3-4d7c8bb2bf2c","added_by":"auto","created_at":"2025-11-24 12:28:12","extension":"png","order_by":37,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":60186,"visible":true,"origin":"","legend":"","description":"","filename":"OnlineFigure3.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/a926255ee0569682bf4e1074.png"},{"id":96629523,"identity":"a0cf1247-8c7f-4db3-8408-98db1608c1b1","added_by":"auto","created_at":"2025-11-24 12:28:12","extension":"png","order_by":38,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":7154,"visible":true,"origin":"","legend":"","description":"","filename":"OnlineFigure4a.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/1a4c17c5064c1450b65fbaea.png"},{"id":96710211,"identity":"0903f812-27cb-4318-a6ae-a82512ca11cc","added_by":"auto","created_at":"2025-11-25 10:10:20","extension":"png","order_by":39,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":8259,"visible":true,"origin":"","legend":"","description":"","filename":"OnlineFigure4b.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/62ed317e3e3841f71112a9c3.png"},{"id":96629529,"identity":"8443e3c5-12b8-4483-b96d-1ff0d250167d","added_by":"auto","created_at":"2025-11-24 12:28:12","extension":"png","order_by":40,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":28210,"visible":true,"origin":"","legend":"","description":"","filename":"OnlineFigure5a.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/d408d823b132e5ea2bb90e6f.png"},{"id":96629531,"identity":"693985de-3f91-4d72-8029-c53e0e31e00c","added_by":"auto","created_at":"2025-11-24 12:28:12","extension":"png","order_by":41,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":26352,"visible":true,"origin":"","legend":"","description":"","filename":"OnlineFigure5b.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/0405bf538decff0528cd0150.png"},{"id":96709033,"identity":"e250aaf1-c8fc-4df0-a4f3-921e980b039c","added_by":"auto","created_at":"2025-11-25 10:07:14","extension":"png","order_by":42,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":25803,"visible":true,"origin":"","legend":"","description":"","filename":"OnlineFigure5c.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/7a0c74ead617d709bbe576dc.png"},{"id":96629540,"identity":"7a2725aa-a3b2-47ef-bfcf-7f476c4d8871","added_by":"auto","created_at":"2025-11-24 12:28:12","extension":"png","order_by":43,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":8908,"visible":true,"origin":"","legend":"","description":"","filename":"OnlineFigure5d.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/3d6b96853f4eea71ee7742e0.png"},{"id":96708497,"identity":"abff0690-e12a-4f5c-94c6-8df73a1e5b08","added_by":"auto","created_at":"2025-11-25 10:03:56","extension":"png","order_by":44,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":23299,"visible":true,"origin":"","legend":"","description":"","filename":"OnlineFigure6a.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/8d87d42a5a5fdcf691375365.png"},{"id":96709911,"identity":"35a5d7f8-e635-46f7-bc3d-f6a38c0d4b5b","added_by":"auto","created_at":"2025-11-25 10:09:47","extension":"png","order_by":45,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":26068,"visible":true,"origin":"","legend":"","description":"","filename":"OnlineFigure6b.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/d152ba67e74a6894cb355b5d.png"},{"id":96629550,"identity":"a55cefad-5240-4b1b-9c78-f562ba2b1c3d","added_by":"auto","created_at":"2025-11-24 12:28:12","extension":"png","order_by":46,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":29944,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/ee8d7069f451183d3429e6af.png"},{"id":96708455,"identity":"7a063f95-9379-406a-aded-1b29f668e562","added_by":"auto","created_at":"2025-11-25 10:02:49","extension":"png","order_by":47,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":158296,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage10.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/ac8f27cbde7a5b86c80f3ae0.png"},{"id":96629530,"identity":"2816df04-821b-4697-a4ff-94f68ca44321","added_by":"auto","created_at":"2025-11-24 12:28:12","extension":"png","order_by":48,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":119490,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/946a541e2aa61e67e0f249f7.png"},{"id":96629533,"identity":"5c05a4a4-44c6-475a-96c2-bd92c87e5a65","added_by":"auto","created_at":"2025-11-24 12:28:12","extension":"png","order_by":49,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":110807,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/ac9ea527110b7b2ac0cbc987.png"},{"id":96629537,"identity":"2594f638-21f7-460c-8d17-92dca8d923e9","added_by":"auto","created_at":"2025-11-24 12:28:12","extension":"png","order_by":50,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":5376,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/bd46bca5dde115ddd1e4e247.png"},{"id":96629535,"identity":"852aafef-8093-4e3c-8c28-e09ad92277aa","added_by":"auto","created_at":"2025-11-24 12:28:12","extension":"png","order_by":51,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":5169,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/7f75611e1699cfcca7bffbf3.png"},{"id":96629545,"identity":"24e09467-0ad2-45b5-8c26-169e1e5e2aa9","added_by":"auto","created_at":"2025-11-24 12:28:12","extension":"png","order_by":52,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":22820,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/de1a4bb87f9fd94845a4d0be.png"},{"id":96710017,"identity":"c2138984-32c2-410b-b273-d495284e674a","added_by":"auto","created_at":"2025-11-25 10:09:53","extension":"png","order_by":53,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":24234,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/2be9c0ee92effc489bf18896.png"},{"id":96629543,"identity":"1c6dc9da-cd54-4cb6-aafe-3ad07a798ac0","added_by":"auto","created_at":"2025-11-24 12:28:12","extension":"png","order_by":54,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":21735,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage8.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/e036b9f4f19deaf2cf73f0d0.png"},{"id":96708587,"identity":"d5763ec7-c008-4fb3-a83c-71b359e030bd","added_by":"auto","created_at":"2025-11-25 10:04:40","extension":"png","order_by":55,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":14798,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage9.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/086de0fb851d172fa16554e8.png"},{"id":96709024,"identity":"728deef6-7233-42e1-8af5-b3300b065c5f","added_by":"auto","created_at":"2025-11-25 10:07:10","extension":"xml","order_by":56,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":136206,"visible":true,"origin":"","legend":"","description":"","filename":"3a8104b9211d4853998fb42ede7887a01structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/411d0c2272f8b13db5bd632e.xml"},{"id":96629548,"identity":"0e9e2c3c-ed51-41eb-bfaa-024eddc39481","added_by":"auto","created_at":"2025-11-24 12:28:12","extension":"html","order_by":57,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":148671,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/7c90ee017987818cdc771152.html"},{"id":96629487,"identity":"657f49e3-7e48-46a4-875a-f63244f0d039","added_by":"auto","created_at":"2025-11-24 12:28:11","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":130587,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eEnd-to-end framework for video-based cattle behavior detection and digital-twin integration.\u003c/strong\u003e Continuous video streams from barn cameras are processed through a pipeline comprising (1) object detection and multi-cow tracking (YOLOv11 + ByteTrack), (2) behavior classification using deep spatiotemporal models (SlowFast / TimeSformer), (3) behavioral data translation into structured activity logs for nutritional modeling via NRC equations, and (4) dynamic synchronization with a Unity-based 3D digital-twin environment for visualization and decision support in precision dairy management\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/d962428f4721afc3893f8bb7.png"},{"id":96629488,"identity":"3d5272f8-1816-46e2-ab78-a4884a13ed45","added_by":"auto","created_at":"2025-11-24 12:28:11","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":276523,"visible":true,"origin":"","legend":"\u003cp\u003eVideo preprocessing and dataset construction workflow. Continuous barn surveillance footage was processed through a multi-stage pipeline involving (1) raw video input from multiple cameras, (2) frame-wise cow detection using YOLOv11, (3) identity-preserving multi-object tracking via ByteTrack, and (4) per-cow bounding box cropping to generate standardized 10-second clips at 224×224 px resolution. This pipeline enabled the creation of a balanced, behavior-focused dataset suitable for spatiotemporal model training and digital-twin integration.\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/bd509d7a5a2d12d8ea4f4f25.png"},{"id":96629490,"identity":"ae417e3f-b953-4bc8-9211-f56c4a03dea2","added_by":"auto","created_at":"2025-11-24 12:28:11","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":515695,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eExamples of spatial and photometric augmentations applied to cow detection frames.\u003c/strong\u003e Augmentation operations included random rotation, cropping, and brightness adjustments to increase visual diversity and improve model robustness to variable camera angles, lighting conditions, and barn environments while preserving ethological interpretability.\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/17900fe92eed0c0f55aa5a79.png"},{"id":96629489,"identity":"919c7828-9273-42c6-b283-50d582897a7f","added_by":"auto","created_at":"2025-11-24 12:28:11","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":33535,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003e\u003cstrong\u003ea \u003c/strong\u003e\u003c/em\u003e\u003cem\u003eNatural behavioral distribution in the un-augmented dataset where lying and standing behaviors dominate while drinking and ruminating are under-represented.\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003e\u003cem\u003e\u003cstrong\u003eb \u003c/strong\u003e\u003c/em\u003e\u003cem\u003eRebalanced distribution after targeted spatial, photometric, and temporal augmentations, which expanded minority classes and improved overall dataset uniformity for model training.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/2df5bc06ae5906026bfc0bc7.png"},{"id":96629496,"identity":"414f69a8-97b2-4bc3-8ace-bcb3898fc23a","added_by":"auto","created_at":"2025-11-24 12:28:11","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":197506,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003e\u003cstrong\u003e(a): \u003c/strong\u003e\u003c/em\u003eSlowFast model trained on the unaugmented dataset.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003e\u003cstrong\u003e(b): \u003c/strong\u003e\u003c/em\u003eSlowFast model trained on the augmented dataset.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003e\u003cstrong\u003e(c): \u003c/strong\u003e\u003c/em\u003eTimeSformer model trained on the unaugmented dataset.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003e\u003cstrong\u003e(d): \u003c/strong\u003e\u003c/em\u003eTimeSformer model trained on the augmented dataset.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003e\u003cstrong\u003e(a-d): Comparative confusion matrices for SlowFast and TimeSformer architectures under unaugmented and augmented training conditions.\u003c/strong\u003e\u003c/em\u003e\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/6125feec7cf4acc46fddbf25.png"},{"id":96629492,"identity":"ca63d645-92a5-42c6-a339-367a19981266","added_by":"auto","created_at":"2025-11-24 12:28:11","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":320428,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003e\u003cstrong\u003eSpatiotemporal attention visualization for representative cattle behaviors\u003c/strong\u003e\u003c/em\u003e\u003cem\u003e. Attention heatmaps generated from the TimeSformer model highlight anatomically and behaviorally relevant regions across frames. During Feeding and Drinking, attention concentrates on the head and muzzle near the feed bunk or water trough, while Lying and Standing states emphasize body and leg contours. In Ruminating behaviors, alternating focus between the head and thoracic areas reflects recognition of cyclical jaw movements. These biologically meaningful attention patterns confirm that the model bases its predictions on ethologically valid visual cues, enhancing interpretability for digital-twin integration and veterinary decision support.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"6.png","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/12dae1c582f4ea56671ad59e.png"},{"id":96712720,"identity":"21081c8a-2597-4c8e-b918-989c782ece87","added_by":"auto","created_at":"2025-11-25 10:16:14","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3291114,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8032374/v1/19608797-eedf-49c4-b738-51d64922ff18.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Video-Based Cattle Behavior Detection for Digital Twin Development in Precision Dairy Systems","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eDigital twin technology represents a paradigm shift in precision livestock farming, transforming how dairy operations monitor, analyze, and optimize animal management through the creation of virtual replicas of physical farm entities \u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, 2\u003c/sup\u003e. These advanced systems integrate real-time sensor data, behavioral patterns, physiological parameters, and environmental conditions to enable continuous monitoring, predictive analytics, and automated decision-making at scales and levels of precision previously unattainable \u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e,\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e. In the dairy sector, digital twins combine behavioral monitoring, mechanistic physiological modeling, nutritional requirement assessment, and environmental impact analysis to generate individualized management strategies that improve both animal welfare and production efficiency \u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e. Empirical evidence supports these benefits: digital twin deployment can enhance feed conversion efficiency by 15\u0026ndash;20%, reduce veterinary interventions by 25\u0026ndash;30%, and cut greenhouse gas emissions per unit of milk by as much as 25% \u003csup\u003e6,2\u003c/sup\u003e. These outcomes illustrate the potential of digital twins as transformative tools for addressing the economic, environmental, and welfare challenges facing modern dairy production. Robust and continuous behavioral monitoring thus serves as the foundation upon which these digital twins operate, providing dynamic, high-frequency data to model individual cow states and management outcomes.\u003c/p\u003e\u003cp\u003eA central prerequisite for digital twin functionality is accurate and continuous behavioral monitoring. In dairy cattle, five behavioral categories like feeding, drinking, lying, standing, and ruminating serve as critical indicators of metabolic balance, reproductive status, disease onset, and welfare status \u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e. Feeding behavior is closely tied to dry matter intake and energy balance, and its measurement is essential for nutritional modeling and ration optimization. Rumination patterns provide direct insights into digestive function and metabolic efficiency, with deviations often preceding metabolic disorders such as acidosis, ketosis, or displaced abomasum \u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e. Lying behavior is a proxy for cow comfort and lameness risk; research has established that 10\u0026ndash;14 hours of daily lying is optimal for both productivity and welfare outcomes\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e. Drinking behavior, while less frequent, is vital for assessing hydration, feed palatability, and heat stress, all of which impact milk yield and welfare\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e. Even standing, often overlooked as a passive activity, can indicate restlessness, discomfort, or estrus expression when measured systematically. Automated recognition of these behaviors is therefore indispensable for creating the continuous, structured data streams that feed into digital twin models \u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e,\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. While the biological relevance of these behavioral metrics is well established, the challenge lies in capturing them continuously and objectively in commercial farm environments.\u003c/p\u003e\u003cp\u003eConventional behavioral monitoring methods, however, present significant limitations that restrict their scalability and accuracy. Manual observation, though capable of capturing fine-grained detail, is labor-intensive, subjective, and economically unsustainable for commercial herds that often number in the hundreds. Even among trained observers, inter-observer variability can reach 15\u0026ndash;25%, reducing the reliability of datasets essential for predictive modeling \u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e,\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e. Wearable sensor systems, while offering continuous monitoring, face a series of practical and welfare-related challenges. Device loss rates of 5\u0026ndash;15% create gaps in longitudinal tracking. Batteries require replacement every 3\u0026ndash;6 months, imposing labor costs and necessitating animal handling that can induce stress. Calibration drift reduces measurement reliability, while sensor placement may alter natural behaviors, potentially biasing the very data intended for welfare and productivity optimization \u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e,\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e. These shortcomings underscore the need for alternative approaches capable of generating high-quality, scalable behavioral data.\u003c/p\u003e\u003cp\u003eComputer vision has emerged as the most promising solution to these limitations, offering a scalable, non-invasive method for monitoring livestock at both individual and group levels. Barn-based video systems enable continuous monitoring of multiple animals simultaneously without interfering with their natural behaviors, while providing richer postural and temporal information than wearable sensors. Video analytics can capture body orientation, head movements, and activity transitions that are critical for distinguishing behaviors such as feeding versus ruminating. Importantly, video-based systems scale cost-effectively, monitoring dozens of animals with a single camera, and integrate easily with existing farm surveillance infrastructure (\u003csup\u003e11\u003c/sup\u003e, \u003csup\u003e2\u003c/sup\u003e). Furthermore, standardized output formats from computer vision pipelines facilitate integration with farm management software, nutritional modeling tools, and digital twin frameworks.\u003c/p\u003e\u003cp\u003eRecent advances in deep learning architectures have dramatically improved the accuracy and robustness of behavior recognition systems. State-of-the-art models have achieved 85\u0026ndash;95% accuracy in multi-class livestock behavior classification tasks \u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e,\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e,\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e. Transformer-based models, originally developed for natural language processing and now adapted for video understanding, show promise for long-duration behaviors. For example, transformer-based classifiers have reached over 90% accuracy in beef cattle behavior recognition, outperforming convolutional networks in capturing spatiotemporal dynamics \u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e. Recent advances in deep learning have improved livestock behavior recognition, with state-of-the-art models reporting 85\u0026ndash;95% accuracy \u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e,\u003csup\u003e9\u003c/sup\u003e,\u003csup\u003e11\u003c/sup\u003e. Transformer-based models show particular promise for long-duration behaviors \u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e. Yet most published systems remain focused on isolated recognition rather than operational digital-twin workflows: identity persistence, structured outputs, and real-time performance under barn variability are rarely addressed. Bridging this gap requires unified pipelines that connect video analytics to nutritional and management models via standardized, digital-twin-ready behavior streams.\u003c/p\u003e\u003cp\u003eSeveral challenges continue to constrain the scalability and robustness of current livestock video-based monitoring systems. Benchmark datasets are typically small (less than 1,000 clips) and collected under controlled conditions, limiting generalization to commercial barns. (\u003csup\u003e10,13\u003c/sup\u003e). The largest publicly available dairy cattle dataset, CBVD-5, contains only 687 clips, insufficient for training deep learning models with robust generalization capability. Moreover, many studies report only offline evaluation results, with real-time performance rarely demonstrated. Yet, for digital twins to function as decision-support tools, latency must remain below 200 ms to provide continuous updates for nutrition, health, and management interventions (\u003csup\u003e1,11\u003c/sup\u003e). Current systems often lack structured outputs, providing only categorical classifications rather than temporally annotated behavioral profiles linked to individual animal IDs. Such structured data streams are critical for downstream integration with nutritional models (e.g., NRC equations), predictive health algorithms, and farm management platforms (\u003csup\u003e4,5\u003c/sup\u003e). Finally, while transformer-based architectures such as TimeSformer have shown strong performance in other domains, their systematic evaluation in livestock contexts remains limited (\u003csup\u003e3,10\u003c/sup\u003e). These gaps highlight the need for comprehensive evaluation frameworks that incorporate real-world datasets, real-time performance, and standardized data outputs tailored to digital twin requirements.\u003c/p\u003e\u003cp\u003eBeyond technical performance, practical deployment factors must also be considered. Barn environments present variable lighting, occlusion, and crowding, all of which complicate computer vision performance. Night-time monitoring often produces reduced accuracy due to low-light conditions, while occlusion from feeding barriers or overlapping animals can obscure key anatomical features necessary for classification. Robust systems must therefore incorporate augmentation strategies and multi-angle camera setups to generalize effectively across such variability. Interpretability is another crucial requirement for industry adoption. Attention mechanisms that highlight anatomically meaningful regions, such as the head for feeding or the legs for lying, provide confidence to farmers, veterinarians, and regulators that the models\u0026rsquo; predictions are biologically valid (\u003csup\u003e7,9\u003c/sup\u003e). Beyond model performance, scalability and cost-effectiveness remain decisive factors for commercial uptake. Systems that function in real time on widely available GPU hardware without excessive computational demands are more likely to achieve practical deployment in dairy barns.\u003c/p\u003e\u003cp\u003eCollectively, these considerations make clear that video-based behavior detection is not only a technical challenge but also a linchpin for the broader adoption of digital twins in dairy farming. By enabling continuous, individualized monitoring, computer vision provides the behavioral data streams required to drive mechanistic nutritional modeling, predictive health analytics, and environmental impact optimization. Without robust and scalable behavior detection systems, the promise of digital twins for livestock remains unattainable.\u003c/p\u003e\u003cp\u003eThis study directly addresses these gaps through the development of a comprehensive video-based cattle behavior detection system explicitly designed for digital twin integration. We present a large-scale dataset collected under authentic barn conditions, incorporating natural environmental variability. An optimized real-time processing pipeline is implemented, combining YOLO detection and ByteTrack tracking for persistent individual cow identification with comparative evaluation of state-of-the-art SlowFast and TimeSformer architectures for behavior classification. Structured outputs, including temporally annotated behavior logs with cow IDs and event durations, are generated in formats compatible with nutritional modeling frameworks and Unity-based visualization systems. By systematically benchmarking performance across models and validating interpretability through spatiotemporal attention analysis, this work establishes a robust technical foundation for integrating behavior detection into dairy digital twins. The objectives of this study are therefore to: (1) construct a large-scale video dataset of core cattle behaviors under commercial barn conditions; (2) develop a real-time multi-animal tracking and classification pipeline; (3) conduct a systematic comparison of SlowFast and TimeSformer architectures; and (4) generate structured, biologically interpretable outputs tailored for digital twin applications in precision dairy systems.\u003c/p\u003e"},{"header":"2. Materials and Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\n \u003ch2\u003e2.1 Experimental Environment and Data Collection\u003c/h2\u003e\n \u003cdiv id=\"Sec4\" class=\"Section3\"\u003e\n \u003ch2\u003e2.1.1 Farm Facility Description\u003c/h2\u003e\n \u003cp\u003eThe study was conducted at the Ruminant Animal Centre (RAC) of Dalhousie University\u0026rsquo;s Agricultural Campus (Truro, Nova Scotia, Canada), a research facility dedicated to advanced dairy production, nutrition, and management studies. Each Holstein cow is tethered in an individual stall with a lying space, a front feed manger and an adjacent in-stall water bucket. Locomotion is constrained to standing and lying within the stall and the cows access feed and water without leaving the stall.\u003c/p\u003e\n \u003cp\u003eEnvironmental control within the barn is achieved through a combination of natural and mechanical ventilation systems. Adjustable sidewall curtains allow air exchange based on external weather conditions, while chimneys with exhaust fans and circulation fans maintain airflow and temperature uniformity. During warmer months, additional ventilation fans are directed toward the cows to alleviate heat stress. The barn follows a 19-hour light and 5-hour dark cycle, with lights turning off at 10:00 p.m. and on at 3:00 a.m. to align with the milking schedule and support circadian rhythm balance.\u003c/p\u003e\n \u003cp\u003eThe RAC herd currently consists of approximately 80 Holstein dairy cows, of which 40 are actively lactating. The animals represent a range of parity and lactation stages, enabling balanced behavioral observations across physiological conditions. Milking is performed twice daily, at 4:30 a.m. and 4:00 p.m., consistent with commercial dairy practices in Atlantic Canada. Cows receive a total mixed ration (TMR) composed of grass silage, corn silage, straw, and concentrate. The formulation is adjusted regularly based on protein and energy analyses of the forage components and the production stage of each group; non-lactating and dry cows receive a lower-energy TMR variant to maintain optimal body condition.\u003c/p\u003e\n \u003cp\u003eAll management and feeding practices comply with Canadian Council on Animal Care (CCAC) guidelines and standard Canadian dairy production protocols. The RAC maintains stringent biosecurity and welfare standards, ensuring that all animals receive routine veterinary oversight, comfortable housing, and nutritional management aligned with NRC (2001) recommendations. All animal procedures were approved by the Dalhousie University Animal Care and Use Committee (Protocol #2024-026, approval date 16‑05‑2024) and complied with CCAC guidelines. An ARRIVE Essential 10 checklist is included in the Supplementary Information.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv id=\"Sec5\" class=\"Section3\"\u003e\n \u003ch2\u003e2.1.2 Multi-Camera Surveillance System\u003c/h2\u003e\n \u003cp\u003eA high-definition closed-circuit surveillance system was installed to enable continuous behavioral monitoring of the dairy herd. The system consisted of a total of seven Panasonic IP cameras, each configured to record at 1920 \u0026times; 1080 resolution with a frame rate of 25\u0026ndash;30 fps.\u003c/p\u003e\n \u003cp\u003eSix of the units were Panasonic WV-S35302-F2L 2MP Outdoor Vandal Dome Cameras, equipped with 2.4 mm fixed lenses, infrared (IR) illumination for low-light monitoring, and integrated microphones for capturing ambient sound. These cameras are IP66 and IK10 rated for environmental durability and impact resistance, and are compliant with FIPS 140-2 Level 3 standards, ensuring secure data handling.\u003c/p\u003e\n \u003cp\u003eOne additional Panasonic WV-X15700-V2L 4K Outdoor Bullet Camera was installed in a high-activity zone. This unit featured a 4.3\u0026ndash;8.6 mm motorized zoom lens and an embedded AI engine capable of supporting up to nine analytic applications simultaneously, enabling high-resolution tracking of fine-grained interactions.\u003c/p\u003e\n \u003cp\u003eCameras were mounted at heights of 3\u0026ndash;4 meters using Panasonic WV-QWL500-W wall brackets and connected via Proterial 61337-8 CAT6A armored plenum-rated cables. Strategic placement at overhead and angled perspectives ensured coverage optimization, overlapping fields of view, and minimization of blind spots, particularly around feed bunks, water troughs, and lying stalls.\u003c/p\u003e\n \u003cp\u003eThe system provided continuous 24/7 monitoring across both day and night cycles, with infrared capabilities enabling uninterrupted observation. To safeguard against data loss, cameras were supported by an UltraTech 1000VA/600W uninterruptible power supply (UPS), ensuring reliability during power fluctuations.\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\n \u003ch2\u003e2.2 Dataset Construction and Annotation Framework\u003c/h2\u003e\n \u003cdiv id=\"Sec7\" class=\"Section3\"\u003e\n \u003ch2\u003e2.2.1 Video Preprocessing Pipeline\u003c/h2\u003e\n \u003cp\u003eThe raw surveillance footage obtained from the multi-camera system was initially subjected to a systematic preprocessing pipeline to ensure its suitability for downstream behavioral analysis. The video streams were first screened manually to identify behaviour rich segments containing clear examples of feeding, drinking, lying and standing. Segments with excessive occlusion, poor lighting or limited behavioural activity were excluded to maintain spatial and temporal consistency across the dataset. This filtering step followed established practices in large-scale livestock video analysis, where minimizing noise is essential for reliable behavioral inference \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e. This pipeline is depicted in Fig. \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e\n \u003cp\u003eFollowing quality filtering, long video sequences were segmented into shorter, behavior-specific clips using a semi-automated workflow. Cows were localized in each frame using a YOLOv11 based object detection model \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e, which was trained on a custom dataset of 2308 images, each labeled with bounding boxes around individual Holstein cows. The dataset was divided into training (1923 images; 83%), validation (193 images; 8%), and test (192 images; 8%) splits. Preprocessing included auto-orientation and extensive data augmentation was applied during training to improve robustness. Augmentation operations \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e included horizontal flips, random resized crops (0\u0026ndash;12% zoom), small rotations between \u0026minus;\u0026thinsp;10\u0026deg; and +\u0026thinsp;10\u0026deg;, brightness adjustments (-15% to +\u0026thinsp;15%), contrast and exposure variations (-10% to +\u0026thinsp;10%), saturation adjustments (-25% to +\u0026thinsp;25%), Gaussian blur (up to 2.5 pixels), and additive noise applied to 0.1% of image pixels. Each training image generated three augmented variants, effectively expanding dataset diversity.\u003c/p\u003e\n \u003cp\u003eThe YOLOv11 model was trained for 90 epochs with a batch size of 16 and a learning rate of 0.01. Training was conducted in Google Colab Pro. The final model achieved high detection performance with mAP@50\u0026thinsp;=\u0026thinsp;0.994, precision\u0026thinsp;=\u0026thinsp;0.982 and recall\u0026thinsp;=\u0026thinsp;0.992, indicating reliable cow detection across varying barn conditions.\u003c/p\u003e\n \u003cp\u003eTo maintain the identity of individual cows across frames and throughout video segments, the detections were linked via the ByteTrack multi-object tracking algorithm \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e, which has been shown to outperform traditional identity-preserving trackers in complex agricultural scenes. This ensured that each cow maintained a consistent ID across frames, even during occlusion or group interactions. Once tracking was established, per-cow bounding boxes were cropped from each frame and resized to 224\u0026times;224 pixels, producing standardized video clips corresponding to individual animals.\u003c/p\u003e\n \u003cp\u003eAll extracted video segments were standardized to a fixed duration of 10 seconds. This length was selected as a compromise between capturing short, transient actions such as drinking, and longer-duration behaviors such as lying or ruminating, which often span several minutes \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e. A fixed temporal window allowed for uniform sampling across the dataset and simplified downstream training procedures. The approach is consistent with recent work in animal behavior recognition, which emphasizes the importance of balancing clip length with the temporal resolution required to distinguish between different behavioral states \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e. As a result, the preprocessing pipeline produced a curated set of behavior-focused clips that maintained both ethological relevance and computational tractability for annotation and model development. The full workflow, from raw video to processed clips, is illustrated in Fig. \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv id=\"Sec8\" class=\"Section3\"\u003e\n \u003ch2\u003e2.2.2 Behavioral Annotation Protocol\u003c/h2\u003e\n \u003cp\u003eBehavioral annotation was performed on the curated 10-second clips to classify cow activity into seven categories: standing and feeding, lying and feeding, drinking, lying, lying and ruminating, standing and ruminating, and standing. These behaviors were selected because they represent the dominant components of the daily activity budget of dairy cattle and are directly linked to productivity, health, and welfare outcomes \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e. Each behavior was defined according to established ethological criteria and verified against visual distinguishability standards to ensure consistent labeling.\u003c/p\u003e\n \u003cp\u003eFeeding was annotated when a cow\u0026rsquo;s head was directed toward the feed bunk, with visible engagement in feed intake \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e. Drinking was assigned when the muzzle was in contact with or directly above the water trough, typically accompanied by head movements consistent with water ingestion \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e. Lying was defined by a resting posture in which the torso was in contact with the stall surface, with limbs folded beneath or alongside the body \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e. Standing was annotated when the cow maintained an upright posture with all four hooves in contact with the ground but without locomotion \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e. Ruminating was identified primarily through cyclical jaw movements associated with cud chewing, which typically occurred during lying but was also observed in stationary standing postures. These operational definitions ensured that categories were mutually exclusive and visually distinguishable in the recorded footage.\u003c/p\u003e\n \u003cp\u003eTo ensure annotation reliability, a dual-annotator protocol was implemented. Two trained annotators independently labeled all clips, and inter-rater agreement was calculated using Cohen\u0026rsquo;s kappa coefficient \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e, which consistently exceeded 0.95 across the dataset, indicating excellent agreement. In cases of disagreement, annotators engaged in consensus discussions to resolve inconsistencies. This multilayered annotation protocol combined ethological rigor, human oversight, and veterinary expertise, ensuring both biological accuracy and reproducibility of the dataset.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv id=\"Sec9\" class=\"Section3\"\u003e\n \u003ch2\u003e2.2.3 Dataset Characteristics\u003c/h2\u003e\n \u003cp\u003eThe final curated dataset comprised 4,964 behavior clips, each standardized to a fixed length of 10 seconds prior to augmentation. The dataset spanned multiple temporal dimensions of variation. Recordings were collected 24/7, thereby capturing fluctuations in behavior associated with environmental changes such as temperature, humidity, and ventilation dynamics. In addition, the dataset encompassed both diurnal cycles and nocturnal activity patterns, supported by infrared camera functionality that enabled continuous monitoring during low-light conditions at night as well. Physiological variation was also represented, as cows at different lactation stages were included, thereby accounting for differences in activity budgets across productive and non-productive phases of the dairy cycle (\u003csup\u003e23,24\u003c/sup\u003e).\u003c/p\u003e\n \u003cp\u003eAnalysis of class distributions revealed a naturally imbalanced dataset, reflecting the time-allocation patterns typical of Holstein dairy cattle. Lying was the most frequent behavior, comprising 24.2% of the dataset, followed by standing (18.3%), feeding (15.4%), ruminating (12.1%), and drinking (3.8%). Such distributions align with established ethological research demonstrating that dairy cows spend a substantial proportion of their daily cycle resting or lying, while drinking occupies only a small fraction of total activity \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\n \u003cp\u003eAlthough the raw dataset exhibited a natural imbalance across behavioral categories, such skewed distributions present a methodological challenge for machine learning, as minority classes like drinking and ruminating may be underrepresented in model training. To address this issue, the imbalance was explicitly corrected in subsequent preprocessing steps through targeted data augmentation (\u003csup\u003e26,27\u003c/sup\u003e). These augmentation strategies expanded the representation of rare behaviors while maintaining the integrity of majority classes, resulting in a more balanced dataset that better supports model generalization. By correcting imbalance at the data preparation stage, the training corpus was aligned both with the ecological diversity of cow behaviors and the computational requirements for robust classification.\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e\n \u003ch2\u003e2.3 Data Augmentation Strategy\u003c/h2\u003e\n \u003cp\u003eTo address the limitations of the naturally imbalanced dataset and to improve the robustness of behavior recognition models, a structured data augmentation strategy was employed. Augmentation has been shown to enhance the generalization of deep learning models by synthetically expanding training data diversity while preserving ethological validity \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e. In this study, augmentation was applied across spatial, photometric, and temporal dimensions, followed by selective class balancing to mitigate underrepresentation of infrequent behaviors.\u003c/p\u003e\n \u003cdiv id=\"Sec11\" class=\"Section3\"\u003e\n \u003ch2\u003e2.3.1 Augmentation Techniques\u003c/h2\u003e\n \u003cp\u003eSpatial augmentations were designed to reduce overfitting to fixed barn layouts and camera viewpoints. Random resized cropping was applied with scaling factors between 0.8 and 1.2, allowing the network to learn from slightly zoomed-in and zoomed-out perspectives. Horizontal flipping was incorporated to mimic mirrored viewpoints, while safe rotations limited to \u0026plusmn;\u0026thinsp;15\u0026deg; introduced natural variability in orientation without compromising behavioral interpretability. These operations helped the model generalize across subtle positional and angular differences that arise from camera placement or cow movement.\u003c/p\u003e\n \u003cp\u003ePhotometric augmentations were introduced to address variability in lighting conditions and sensor noise. Brightness adjustments of \u0026plusmn;\u0026thinsp;20% and contrast modifications of \u0026plusmn;\u0026thinsp;15% simulated the natural fluctuations in illumination across day-night cycles and seasonal changes. Additionally, Gaussian noise with \u0026sigma;\u0026thinsp;=\u0026thinsp;0.02 was added to replicate the visual distortions that occur under low-light or high-contrast recording conditions. By incorporating these variations, the model was trained to remain invariant to non-behavioral visual artifacts while focusing on essential cues for classification.\u003c/p\u003e\n \u003cp\u003eExamples of spatial and photometric augmentations applied to cow images are shown in Fig. \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e. These include random cropping, rotations, brightness adjustments, and noise addition, which increase dataset diversity while preserving ethological interpretability.\u003c/p\u003e\n \u003cp\u003eTemporal augmentations accounted for the inherent variability in behavioral tempo and duration. Frame sampling rates were varied to simulate different effective frame rates, thereby ensuring that the model learned representations that were robust to temporal resolution changes. In addition, clip duration modifications were performed within a safe margin around the standard 10-second window, allowing the model to handle both shorter clips (e.g., drinking events) and slightly extended sequences (e.g., lying or ruminating). These operations encouraged the network to capture temporal dynamics of behavior at multiple granularities, a practice consistent with best practices in spatio-temporal modeling \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e28\u003c/span\u003e,\u003cspan class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv id=\"Sec12\" class=\"Section3\"\u003e\n \u003ch2\u003e2.3.2 Class Balancing Implementation\u003c/h2\u003e\n \u003cp\u003eIn addition to enhancing data diversity, augmentation was selectively employed to address the class imbalance observed in the raw dataset. Rare behaviors such as drinking and ruminating were deliberately oversampled through targeted augmentation factors, while more frequent behaviors were augmented conservatively. Specifically, drinking clips were augmented 7.5-fold, ruminating clips 4.5-fold, and the remaining behaviors (feeding, standing, lying) by 2-3-fold.\u003c/p\u003e\n \u003cp\u003eThis strategy expanded the dataset to over 9,600 clips, resulting in a substantially more balanced distribution across the behaviors. The final dataset composition included feeding (22.9%), standing (21.9%), lying (20.8%), ruminating (18.8%), and drinking (15.6%). By mitigating imbalance while preserving the ecological realism of behavior patterns, the augmented dataset provided a robust foundation for training behavior recognition models capable of performing reliably in real-world farm environments. The effect of the augmentation pipeline on class balance is illustrated in Fig. 4a and Fig. 4b, which compares the behavioral class distributions of the unaugmented and augmented datasets. As shown, augmentation substantially increased the representation of minority behaviors such as drinking and ruminating thereby improving dataset uniformity for model training.\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec13\" class=\"Section2\"\u003e\n \u003ch2\u003e2.4 Deep Learning Architecture Implementation\u003c/h2\u003e\n \u003cp\u003eTo evaluate spatiotemporal modeling approaches for cow behavior recognition, two state-of-the-art video classification architectures were implemented: the SlowFast network \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e and the TimeSformer model \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e. These architectures were selected because they represent complementary strategies for capturing both short-term motion dynamics and long-range temporal dependencies, which are essential for distinguishing between rapid actions such as drinking and extended behaviors such as lying or ruminating.\u003c/p\u003e\n \u003cdiv id=\"Sec14\" class=\"Section3\"\u003e\n \u003ch2\u003e2.4.1 SlowFast Network Configuration\u003c/h2\u003e\n \u003cp\u003eThe SlowFast network employs a dual-pathway design, consisting of a slow pathway that processes frames at a low temporal resolution and a fast pathway that captures finer temporal granularity. In this study, the slow pathway operated on 8 frames per clip sampled at fixed intervals, while the fast pathway processed 32 frames per clip at a reduced spatial resolution of 224\u0026times;224 pixels. This configuration allowed the slow pathway to capture global semantic context such as posture (e.g., lying vs. standing), while the fast pathway focused on short-term dynamics such as head movements during feeding or drinking.\u003c/p\u003e\n \u003cp\u003eThe two pathways were connected via lateral fusion layers, enabling information flow between slow and fast branches. These connections ensured that fine-grained motion features extracted by the fast branch enriched the semantic representations of the slow branch. Both pathways used a 3D convolutional backbone (ResNet-style), with spatiotemporal kernels applied across the frame sequence to jointly model motion and appearance. Computational efficiency was achieved by applying fewer channels in the fast pathway (1/8 of the slow pathway), reducing redundancy while preserving motion sensitivity. This design is particularly suited to livestock behavior analysis, as it mirrors the temporal scales of dairy cow activity budgets: fast dynamics (head/muzzle motion during feeding or drinking) and slow dynamics (postural changes such as lying or standing) \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv id=\"Sec15\" class=\"Section3\"\u003e\n \u003ch2\u003e2.4.2 TimeSformer Architecture Design\u003c/h2\u003e\n \u003cp\u003eThe TimeSformer model adopts a transformer-based architecture that replaces 3D convolutions with divided spatiotemporal self-attention mechanisms. Each input video clip was decomposed into non-overlapping patches, which were linearly embedded and fed into transformer encoder layers. Attention was computed separately along spatial and temporal dimensions, reducing computational cost compared to full joint attention while retaining modeling power \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\n \u003cp\u003eIn the spatial attention module, multi-head self-attention with 8 heads was applied to capture dependencies among different regions within each frame. In parallel, the temporal attention module modeled relationships across frames, allowing the network to capture long-range dependencies such as the transition from standing to lying or sustained ruminating bouts. To ensure position awareness, learned positional encodings were incorporated, enabling the model to distinguish between patches based on both spatial location and temporal order.\u003c/p\u003e\n \u003cp\u003eA hybrid attention scheme was employed, in which layers alternated between spatial and temporal attention. This design provided the model with a global receptive field across both space and time, while maintaining computational tractability. By leveraging self-attention, the TimeSformer was able to integrate contextual information across the entire clip, making it especially effective for recognizing behaviors characterized by subtle, distributed cues.\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec16\" class=\"Section2\"\u003e\n \u003ch2\u003e2.5 Training Configuration and Optimization\u003c/h2\u003e\n \u003cdiv id=\"Sec17\" class=\"Section3\"\u003e\n \u003ch2\u003e2.5.1 Hardware and Software Setup\u003c/h2\u003e\n \u003cp\u003eAll experiments were conducted on a cloud based GPU platform (Google LLC, USA), which provided access to an NVIDIA L4 GPU (24 GB VRAM) for model training. This cloud-based setup ensured sufficient memory and computational capacity for handling spatiotemporal video architectures such as 3D CNNs and transformers. Local preprocessing, annotation handling, and lightweight experiments were performed on a MacBook Air with Apple M3 chip (8-core CPU, integrated GPU, 16 GB unified memory).\u003c/p\u003e\n \u003cp\u003eThe computational environment was configured on Ubuntu 20.04 LTS with CUDA 11.6 and PyTorch 1.12.0 serving as the primary deep learning framework. The environment is also equipped with standard libraries for computer vision and video processing, including OpenCV 4.7, NumPy, and scikit-learn. This setup provided a reproducible and flexible platform for both development and large-scale training.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv id=\"Sec18\" class=\"Section3\"\u003e\n \u003ch2\u003e2.5.2 Model-Specific Training Parameters\u003c/h2\u003e\n \u003cp\u003eModel training was tailored to the architectural and computational requirements of each network. For the SlowFast network \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e, training was performed with a batch size of 4, using the Adam optimizer with a learning rate of 1 \u0026times; 10⁻⁴. The loss function was cross-entropy, and the model was trained for 10 epochs, with extensions to 20 when necessary. Data loading employed two worker threads (num_workers\u0026thinsp;=\u0026thinsp;2), with shuffling enabled to maximize diversity.\u003c/p\u003e\n \u003cp\u003eFor the TimeSformer model \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e, a more advanced optimization scheme was required to stabilize transformer training. Each input consisted of 12 frames of 224\u0026times;224 resolution, processed in batches of 4 samples, with gradient accumulation across 4 steps to achieve an effective batch size of 16. Training proceeded for 20 epochs using the AdamW optimizer, with a learning rate of 3 \u0026times; 10⁻⁵, weight decay of 0.01, and a warmup ratio of 0.1 to gradually ramp learning in the early iterations. To ensure numerical stability, gradients were clipped at a norm of 1.0, and mixed precision training was enabled to improve efficiency. Early stopping was applied with a patience of 4 epochs, halting training when validation performance failed to improve.\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec19\" class=\"Section2\"\u003e\n \u003ch2\u003e2.6 Real-Time Processing Pipeline\u003c/h2\u003e\n \u003cdiv id=\"Sec20\" class=\"Section3\"\u003e\n \u003ch2\u003e2.6.1 End-to-End System Architecture\u003c/h2\u003e\n \u003cp\u003eThe proposed framework was implemented as a modular end-to-end video analysis pipeline, enabling automated behavior recognition from continuous surveillance footage of dairy cows. The system integrated three major components: object detection and tracking, behavior classification, and temporal smoothing.\u003c/p\u003e\n \u003cp\u003eRaw video input from the multi-camera surveillance system was first processed through the YOLOv11 object detection model, trained on manually annotated cow images as described in Section 4.2.1. This stage localized individual cows within each frame. The resulting detections were then passed to the ByteTrack multi-object tracking algorithm, which maintained consistent cow identities across successive frames and prevented errors due to occlusions or group interactions. This ensured that each animal\u0026rsquo;s behavior was modeled continuously over time rather than as isolated instances.\u003c/p\u003e\n \u003cp\u003eThe cropped, per-cow video clips were subsequently fed into the behavior classification models (SlowFast or TimeSformer), which predicted one of seven behavioral states: standing and feeding, drinking, lying and feeding, lying, standing, standing and ruminating or lying and ruminating. To accommodate model input requirements, a sliding window strategy was employed, where clips of 16\u0026ndash;32 consecutive frames were extracted with a 50% overlap and an 8-frame stride. This design captured sufficient temporal context for behavior recognition while ensuring efficient use of computational resources.\u003c/p\u003e\n \u003cp\u003eFinally, to improve robustness against frame-level misclassifications, the system incorporated majority voting and temporal consistency enforcement. Predicted labels within overlapping windows were aggregated, and transitions were only accepted if they persisted across multiple consecutive windows. This reduced spurious label switching and produced smoother, ethologically plausible behavior timelines. By combining real-time detection, identity-preserving tracking, spatiotemporal classification, and temporal smoothing, the pipeline provided a reliable end-to-end solution for continuous monitoring of cow behavior in commercial farm environments.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv id=\"Sec21\" class=\"Section3\"\u003e\n \u003ch2\u003e2.6.2 Output Generation and Format Standards\u003c/h2\u003e\n \u003cp\u003eThe final stage of the pipeline was responsible for converting model predictions into structured outputs suitable for downstream analysis, integration with nutritional modeling frameworks, and visualization in digital twin environments. To ensure both interoperability and reproducibility, standardized formats were adopted for data storage and streaming.\u003c/p\u003e\n \u003cp\u003eAll predictions were logged into structured CSV files containing the following fields: cow ID, predicted activity label, start timestamp, end timestamp, activity duration, and model confidence score. These logs enabled quantitative evaluation of behavioral activity budgets while preserving the temporal context of each prediction. Each entry was aligned with the synchronized video timeline, ensuring that results could be directly traced back to the original raw footage.\u003c/p\u003e\n \u003cp\u003eTo support domain-specific applications, the output schema was designed for compatibility with NRC-based nutritional modeling frameworks, where activity budgets (e.g., feeding or lying duration) inform energy expenditure and intake calculations. In addition, the logs were formatted for seamless integration with Unity-based visualization systems, where predicted states were used to drive cow avatars in a digital twin of the barn environment. This facilitated intuitive, real-time interpretation of behavioral patterns by both researchers and farm managers.\u003c/p\u003e\n \u003cp\u003eFor real-time applications, the system also incorporated low-latency streaming capabilities. Predictions were generated and transmitted in near real-time, with latency optimization achieved through GPU-accelerated inference, sliding window buffering, and asynchronous log writing. This ensured that activity updates were available within seconds of observation, enabling the framework to function not only as a retrospective analysis tool but also as a foundation for real-time monitoring and decision support systems in precision dairy farming.\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e"},{"header":"3. Results","content":"\u003cdiv id=\"Sec23\" class=\"Section2\"\u003e\n \u003ch2\u003e3.1 Dataset Characterization and Distribution Analysis\u003c/h2\u003e\n \u003cdiv id=\"Sec24\" class=\"Section3\"\u003e\n \u003ch2\u003e3.1.1 Pre- and Post-Augmentation Comparison\u003c/h2\u003e\n \u003cp\u003eThe annotated dataset comprised seven distinct behavioral classes extracted from continuous barn recordings. As summarized in Fig.\u0026nbsp;4, the augmented dataset achieved a near-uniform class balance following the targeted oversampling strategy described in Section 4.3.2.\u003c/p\u003e\n \u003cp\u003ePost-augmentation, the behavioral proportions aligned closely with known ethological activity budgets of Holstein cows lying and standing remained the most frequent states, while drinking and ruminating increased to biologically realistic levels. This balanced representation provided a stable foundation for downstream model training, ensuring that minority behaviors were adequately represented without distorting the natural behavioral hierarchy.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv id=\"Sec25\" class=\"Section3\"\u003e\n \u003ch2\u003e3.1.2 Natural Behavioral Frequency Analysis\u003c/h2\u003e\n \u003cp\u003eThe overall behavioral distribution closely aligns with known diurnal activity patterns of Holstein cows under tie-stall housing conditions. Lying and standing behaviors were predominant during non-feeding hours, reflecting periods of rest and comfort \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e. Feeding activity peaked during scheduled feed deliveries (typically morning and late afternoon), while ruminating bouts were interspersed throughout the day, particularly following feeding episodes \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e23\u003c/span\u003e\u003c/sup\u003e. Drinking events occurred less frequently but were distributed evenly across the photoperiod \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e32\u003c/span\u003e\u003c/sup\u003e. These observations confirm that the collected dataset reflects biologically realistic behavioral proportions rather than sampling artifacts.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv id=\"Sec26\" class=\"Section3\"\u003e\n \u003ch2\u003e3.1.3 Augmentation Effectiveness Assessment\u003c/h2\u003e\n \u003cp\u003eThe augmentation pipeline effectively expanded the training set size while preserving spatiotemporal consistency and behavioral realism. Augmented samples retained anatomical fidelity (e.g., head alignment with feed bunks or water troughs) and contextual cues (stall boundaries, lighting gradients) critical for accurate model training. Quantitatively, class balance improved from an initial max:min ratio of ~\u0026thinsp;9:1 to approximately 2:1, reducing overfitting risk and improving minority-class recognition during preliminary validation runs.\u003c/p\u003e\n \u003cp\u003eCollectively, these preprocessing and augmentation procedures produced a balanced, ecologically valid dataset capable of supporting robust training of deep learning models for fine-grained cow behavior recognition.\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec27\" class=\"Section2\"\u003e\n \u003ch2\u003e3.2 Model Performance Comparative Analysis\u003c/h2\u003e\n \u003cp\u003eThe comparative evaluation of TimeSformer and SlowFast models aimed to determine the optimal spatiotemporal architecture for recognizing complex cattle behaviors under realistic barn conditions. Both models were trained and validated on the same curated seven-class dataset encompassing Drinking, Feeding and Lying, Feeding and Standing, Lying, Ruminating and Lying, Ruminating and Standing, and Standing. Each model was trained for 10 epochs using Adam optimization (learning rate\u0026thinsp;=\u0026thinsp;1 \u0026times; 10⁻⁴, batch size\u0026thinsp;=\u0026thinsp;8) and identical augmentation pipelines (horizontal flip, temporal jittering, illumination normalization). The dataset was split into 70% training, 15% validation, and 15% testing sets.\u003c/p\u003e\n \u003cdiv id=\"Sec28\" class=\"Section3\"\u003e\n \u003ch2\u003e3.2.1 Overall Classification Accuracy\u003c/h2\u003e\n \u003cp\u003eAcross all evaluation runs, the TimeSformer achieved a mean overall accuracy of 85.0%, exceeding the SlowFast baseline accuracy of 82.3%. This improvement reflects the transformer\u0026rsquo;s inherent strength in capturing long-range temporal dependencies via its divided space-time self-attention mechanism. By encoding each video clip as a series of patch-level tokens, TimeSformer effectively attends to extended motion sequences such as continuous feeding or rumination cycles. The SlowFast model, although powerful in detecting high-frequency actions like drinking or head-lifting, relies primarily on convolutional hierarchies that emphasize localized motion and hence can miss broader behavioral context when frame-to-frame differences are minimal.\u003c/p\u003e\n \u003cp\u003eTo ensure the robustness of the observed performance gap, five independent trials were executed using different random initializations. The mean accuracy difference of 2.7 percentage points was statistically significant (p\u0026thinsp;=\u0026thinsp;0.031, \u0026alpha;\u0026thinsp;=\u0026thinsp;0.05) based on a paired-sample t-test. This reproducibility underscores the stability of the transformer encoder for real-world deployment, where models must maintain reliable behavior detection despite day-to-day environmental variability. The superior accuracy of TimeSformer therefore validates the hypothesis that attention-driven global context modeling is essential for continuous barn surveillance, in which many cow behaviors manifest over extended time horizons rather than through abrupt motion cues alone.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv id=\"Sec29\" class=\"Section3\"\u003e\n \u003ch2\u003e3.2.2 Per-Class Performance Evaluation\u003c/h2\u003e\n \u003cp\u003eA fine-grained comparison of precision, recall, and F1-scores () provides further insight into class-specific performance trends. Table \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e summarizes the observed metrics and the corresponding improvements (\u0026Delta;F1) obtained by the transformer-based model.\u003c/p\u003e\n \u003cdiv class=\"gridtable\"\u003e\n \u003ctable id=\"Tab1\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003e\u003cstrong\u003ePer-class F1-scores for SlowFast and TimeSformer models across seven cattle behavior categories.\u003c/strong\u003e The TimeSformer consistently outperformed the SlowFast network for most behaviors, particularly in Feeding \u0026amp; Standing, Standing, and Ruminating \u0026amp; Standing, reflecting its superior capacity to capture long-range temporal dependencies. Minor reductions for static postures (Lying and Ruminating \u0026amp; Lying) suggest that convolution-based architectures remain advantageous for low-motion states.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eBehavior Class\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eSlowFast F1\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eTimeSformer F1\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eDrinking\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.910\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.901\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eFeeding \u0026amp; Lying\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.854\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.831\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eFeeding \u0026amp; Standing\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.924\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.947\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eLying\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.888\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.878\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eRuminating \u0026amp; Lying\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.806\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.786\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eRuminating \u0026amp; Standing\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.742\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.786\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eStanding\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.667\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.759\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eMacro Average\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.827\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.861\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eWeighted Average\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.851\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.861\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003cp\u003eThe largest improvement was recorded for Standing followed by Ruminating \u0026amp; Standing. These behaviors exhibit limited limb motion but prolonged duration, where the attention mechanism excels at integrating subtle temporal features - such as slight head tilts or jaw motions spread over many frames. Conversely, modest F1-score reductions for static postures (Lying, Ruminating \u0026amp; Lying) suggest that when motion cues are almost absent, TimeSformer\u0026rsquo;s patch-level embeddings may average out fine-grained micro-movements.\u003c/p\u003e\n \u003cp\u003eMacro-averaged F1 improved from 0.827 (SlowFast) to 0.841 (TimeSformer), while the weighted F1 increased from 0.851 to 0.861. These gains confirm that transformer models generalize better across majority and minority behavior categories, yielding a more balanced classification under heterogeneous barn environments. Moreover, TimeSformer\u0026rsquo;s superior performance under nighttime and occluded conditions demonstrates its ability to learn illumination-invariant spatiotemporal representations, a crucial requirement for continuous, unattended surveillance in digital-twin pipelines.\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec30\" class=\"Section2\"\u003e\n \u003ch2\u003e3.3 Confusion Matrix and Error Analysis\u003c/h2\u003e\n \u003cp\u003eWhereas overall and per-class metrics quantify accuracy, confusion-matrix analysis exposes systematic misclassification patterns and the behavioral semantics behind them.\u003c/p\u003e\n \u003cp\u003eTo examine class-wise prediction behavior under different training conditions, confusion matrices were generated for both unaugmented and augmented datasets for each architecture. Figures 5(a-d) present a side-by-side comparison: (a) SlowFast - Unaugmented, (b) SlowFast - Augmented, (c) TimeSformer - Unaugmented, and (d) TimeSformer - Augmented. All matrices use identical color scaling for direct comparability.\u003c/p\u003e\n \u003cdiv id=\"Sec31\" class=\"Section3\"\u003e\n \u003ch2\u003e3.3.1 Side-by-Side Confusion Matrix Comparison\u003c/h2\u003e\n \u003cp\u003eBoth matrices exhibit strong diagonal dominance, confirming that most behaviors were correctly identified. However, persistent confusion clusters reveal key perceptual ambiguities inherent in cattle behavior recognition:\u003c/p\u003e\n \u003cul\u003e\n \u003cli\u003e\n \u003cp\u003eLying vs. Ruminating \u0026amp; Lying (~\u0026thinsp;12\u0026ndash;15% cross-confusion): These categories differ mainly through subtle jaw-movement cues. SlowFast\u0026rsquo;s motion-centric filters sometimes failed to detect small cyclic mandibular motions, labeling ruminating cows as merely lying. TimeSformer\u0026rsquo;s attention mechanism partially alleviated this overlap by tracking head-region temporal changes, thereby reducing off-diagonal errors.\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eStanding vs. Ruminating \u0026amp; Standing: Both states share identical postural geometry, making visual separation challenging. TimeSformer achieved a higher \u003cem\u003eStanding\u003c/em\u003e F1 (0.75 vs. 0.66) due to its capacity to integrate weak but consistent temporal signals, such as rhythmic chewing motions persisting across several seconds.\u003c/p\u003e\n \u003c/li\u003e\n \u003c/ul\u003e\n \u003cp\u003eOverall, the TimeSformer confusion matrix exhibits higher diagonal purity and reduced class-spillover, underscoring its superior ability to encode fine-grained temporal semantics. In practical terms, these improvements translate directly into more precise behavioral state transitions within the digital twin, ensuring that the simulated cow entities maintain realistic time-aligned activity cycles.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv id=\"Sec32\" class=\"Section3\"\u003e\n \u003ch2\u003e3.3.2 Error Source Attribution and Visual Inspection\u003c/h2\u003e\n \u003cp\u003eA post-hoc qualitative review of misclassified clips was performed to trace error origins and assess their implications for digital-twin fidelity.\u003c/p\u003e\n \u003cp\u003eThree dominant sources were identified:\u003c/p\u003e\n \u003col\u003e\n \u003cli\u003eBehavioral Ambiguity: Many errors stem from the continuum between similar activities. Static postures punctuated by micro-movements often blur class boundaries. For example, a cow transitioning from Lying to Ruminating \u0026amp; Lying may perform slow, imperceptible head adjustments that span multiple frames-challenging both convolutional and attention-based encoders. These findings emphasize the biological fluidity of behavior categories, suggesting that discrete classification may need to evolve toward probabilistic or continuous state modeling in future digital-twin iterations.\u003c/li\u003e\n \u003cli\u003eEnvironmental and Lighting Variability: The barn environment introduced dynamic illumination, shadowing, and occlusion events. During dusk or under mixed natural-artificial lighting, contrast reduction led to occasional false positives in Drinking when reflective surfaces mimicked mouth-to-trough contact. Although TimeSformer exhibited stronger illumination resilience, both models showed mild accuracy drops in nocturnal segments, implying that domain-adaptive normalization or temporal brightness augmentation could further enhance generalization.\u003c/li\u003e\n \u003cli\u003eModel Representation Bias: Minority classes (Feeding \u0026amp; Lying, Ruminating \u0026amp; Standing) contained fewer labeled instances, yielding under-optimized feature representations. The transformer\u0026rsquo;s tokenization may dilute critical motion vectors, while SlowFast\u0026rsquo;s motion-intensity bias sometimes over-fitted to fast activities.\u003c/li\u003e\n \u003c/ol\u003e\n \u003cp\u003eMitigation strategies include focal loss weighting, synthetic clip generation, and attention-map regularization to improve the salience of under-represented actions.\u003c/p\u003e\n \u003cp\u003eImplications for Digital-Twin Deployment: Misclassifications manifest as short-term state errors in the virtual twin, occasionally causing temporal discontinuities or overcounting of specific behaviors (e.g., fragmented rumination episodes). However, incorporating temporal smoothing, majority-vote state correction, and behavior-duration thresholds effectively attenuates these artifacts, preserving realistic simulation continuity. The qualitative review thus reinforces the importance of explainable model diagnostics-through attention-map visualization and per-frame attribution-to support transparent, trustworthy integration of computer-vision outputs into precision-nutrition digital twins.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv id=\"Sec33\" class=\"Section3\"\u003e\n \u003ch2\u003e3.3.3 Architecture-Specific Strengths and Weaknesses\u003c/h2\u003e\n \u003cp\u003eA comparative assessment of both models reveals complementary strengths arising from their distinct spatiotemporal encoding strategies.\u003c/p\u003e\n \u003cp\u003eThe TimeSformer, leveraging transformer-based self-attention, demonstrates superior sensitivity to feeding and drinking behaviors, where subtle head and mouth motions dominate. Its patch-token attention distributes weights dynamically across spatial regions, enabling the model to integrate micro-movements within longer temporal windows. Consequently, TimeSformer effectively identifies sequences where feeding occurs intermittently amid minimal body displacement, an ability crucial for capturing nutritional intake events in the digital twin.\u003c/p\u003e\n \u003cp\u003eConversely, the SlowFast architecture excels at posture driven behaviors such as Lying, Standing, and Ruminating \u0026amp; Lying. Its dual-pathway design where the fast branch encodes fine temporal granularity and the slow branch captures broader spatial context allows strong geometric posture recognition even under low motion conditions. The network\u0026rsquo;s convolutional inductive biases help stabilize predictions in cluttered environments, reducing false positives for large-scale static postures.\u003c/p\u003e\n \u003cp\u003eError-pattern inspection further reveals biologically consistent tendencies: misclassifications typically occur between physiologically adjacent states, such as transitions between Ruminating \u0026amp; Standing and Standing, or Lying and Ruminating \u0026amp; Lying. These correspond to authentic behavioral continuums rather than algorithmic artifacts, confirming that both models mirror the natural fluidity of bovine activities rather than imposing arbitrary separations. Hence, the observed confusions hold biological interpretability, underscoring the potential of deep spatiotemporal models as quantitative tools for ethological research.\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec34\" class=\"Section2\"\u003e\n \u003ch2\u003e3.4 Spatiotemporal Attention Visualization\u003c/h2\u003e\n \u003cdiv id=\"Sec35\" class=\"Section3\"\u003e\n \u003ch2\u003e3.4.1 Attention Heatmap Analysis\u003c/h2\u003e\n \u003cp\u003eTo interpret model decision processes, attention-map visualizations \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e33\u003c/span\u003e,\u003cspan class=\"CitationRef\"\u003e34\u003c/span\u003e\u003c/sup\u003e were generated for representative clips of each behavioral class. These visualizations reveal that learned focus areas align with anatomically relevant regions, reinforcing biological validity.\u003c/p\u003e\n \u003cp\u003eDuring Feeding and Drinking, approximately 65% of cumulative attention weight concentrated around the head and muzzle region, particularly near the feed trough and water bucket interfaces. This indicates that the transformer successfully prioritized regions corresponding to ingestion activity, capturing the periodic forward-backward jaw motion characteristic of feeding bouts.\u003c/p\u003e\n \u003cp\u003eFor Standing and Lying postures, attention activation shifted toward body and leg contours, representing the model\u0026rsquo;s recognition of skeletal orientation and support distribution. In Ruminating \u0026amp; Lying states, attention maps alternated between head and thoracic regions, reflecting internalized understanding of chewing cycles within stationary frames.\u003c/p\u003e\n \u003cp\u003eThe corresponding attention visualizations are presented in Fig. \u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e, which illustrate the spatial concentration of the model\u0026rsquo;s focus across representative behavioral classes. As shown, the TimeSformer consistently directs attention toward anatomically relevant regions such as the head and muzzle during feeding and drinking or the torso and limbs during postural states demonstrating biologically coherent feature learning. These spatial distributions confirm that model activations correspond to meaningful anatomical cues rather than spurious background patterns. By validating attention heatmaps against known ethological indicators-such as head movement frequency and body-weight distribution-the study establishes a direct link between learned visual salience and biological interpretability.\u003c/p\u003e\n \u003cp\u003eThis level of interpretability provides actionable transparency for veterinary decision support. By visualizing which regions and temporal segments drive classification, farm operators can corroborate algorithmic outputs with on-ground observations, improving trust in automated monitoring systems. In the context of the digital twin, attention transparency facilitates the translation of model activations into behavioral biomarkers, supporting nutritional adjustment decisions, welfare diagnostics, and anomaly detection.\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec36\" class=\"Section2\"\u003e\n \u003ch2\u003e3.5 Real-Time Processing Performance\u003c/h2\u003e\n \u003cdiv id=\"Sec37\" class=\"Section3\"\u003e\n \u003ch2\u003e3.5.1 Computational Efficiency Benchmarking\u003c/h2\u003e\n \u003cp\u003eBoth architectures were benchmarked on an NVIDIA RTX A100 GPU on google colab using mixed-precision inference. The TimeSformer achieved an average throughput of 22.6 fps, surpassing SlowFast\u0026rsquo;s 18.4 fps due to its efficient token-wise computation and reduced temporal redundancy. Despite its transformer complexity, TimeSformer maintained lower GPU memory consumption (7.2 GB) compared to SlowFast\u0026rsquo;s 8.9 GB, attributed to the absence of multi-branch convolutional streams.\u003c/p\u003e\n \u003cp\u003eFrom a deployment perspective, both models meet real-time processing thresholds for 25 fps video feeds \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e35\u003c/span\u003e\u003c/sup\u003e. However, the TimeSformer\u0026rsquo;s higher parallelization efficiency and consistent batch-to-batch latency (~\u0026thinsp;44 ms/frame) make it preferable for continuous inference within digital-twin pipelines where computational scalability and stability are essential. A comparative hardware analysis indicates that real-time deployment is feasible on high-end consumer GPUs or edge-AI modules (e.g., Jetson AGX Orin) with slight down-sampling. Such resource profiling ensures that behavior recognition remains sustainable for on-farm automation without excessive power draw.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv id=\"Sec38\" class=\"Section3\"\u003e\n \u003ch2\u003e3.5.2 Scalability and Multi-Cow Processing\u003c/h2\u003e\n \u003cp\u003eTo emulate barn-scale conditions, the end-to-end system was evaluated on simultaneous video streams tracking 16 individual cows using YOLOv11\u0026thinsp;+\u0026thinsp;ByteTrack for detection and identity assignment. The combined detection-tracking-classification pipeline achieved an average latency of 180 ms per frame, corresponding to real-time throughput at 5.5 fps per cow.\u003c/p\u003e\n \u003cp\u003eBatch-processing optimization and asynchronous GPU queues reduced total inference time by 27%, confirming the framework\u0026rsquo;s scalability to group monitoring without significant degradation in accuracy. These benchmarks demonstrate the system\u0026rsquo;s capacity for large-herd digital-twin synchronization, where multiple physical cows can be updated concurrently in the virtual environment. Future work will explore multi-GPU distribution and temporal batching to achieve near-real-time herd-level behavior analytics.\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec39\" class=\"Section2\"\u003e\n \u003ch2\u003e3.6 Challenge Analysis and Performance Limitations\u003c/h2\u003e\n \u003cdiv id=\"Sec40\" class=\"Section3\"\u003e\n \u003ch2\u003e3.6.1 Environmental Variability Impact\u003c/h2\u003e\n \u003cp\u003eDespite strong baseline accuracy, environmental factors continued to influence detection reliability. Under night-time illumination, accuracy declined by 8\u0026ndash;12%, primarily due to reduced contrast and color-channel noise in infrared footage \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e. Occlusions-caused by barn structures, equipment, or cow overlapped to intermittent trajectory loss, occasionally truncating behavior sequences \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. The tracking system successfully reidentified 92% of interrupted tracks, but brief identity swaps introduced minor temporal labeling noise. Corner and edge regions of the camera field exhibited degraded detection consistency, as cows partially exited the frame, limiting spatial continuity for both TimeSformer and SlowFast. Incorporating multi-view camera fusion and adaptive brightness equalization could mitigate these edge-case degradations.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv id=\"Sec41\" class=\"Section3\"\u003e\n \u003ch2\u003e3.6.2 Behavior-Specific Recognition Challenges\u003c/h2\u003e\n \u003cp\u003ePersistent confusion between Lying and Ruminating \u0026amp; Lying remains the principal classification bottleneck. Both behaviors share near-identical postural geometry; differentiation relies solely on subtle mandibular motion, which is occasionally obscured or temporally aliased at low frame rates \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e36\u003c/span\u003e\u003c/sup\u003e. Enhancing temporal resolution or integrating optical-flow-based motion cues could improve distinction in future iterations. Drinking behavior detection also posed challenges due to reflective water surfaces and variable head angles. False positives occurred when cows lowered their heads near troughs without actual ingestion, emphasizing the need for multi-modal fusion with acoustic or RFID-based water-intake sensors.\u003c/p\u003e\n \u003cp\u003eFinally, temporal-sequence modeling limitations were evident in transitions between behaviors. Both architectures occasionally produced fragmented predictions during behavior shifts (e.g., standing to lying), suggesting that explicit temporal smoothing or recurrent attention integration could enhance continuity. These findings emphasize that while current deep-video architectures excel at single-state classification, the goal of continuous, biologically faithful behavioral tracking requires hybrid temporal models combining deep attention with probabilistic state transition logic.\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e"},{"header":"4. Discussion","content":"\u003cp\u003eThe TimeSformer architecture demonstrated distinct advantages for livestock video analytics owing to its global spatiotemporal attention mechanism, which captures both spatial structures and long-term temporal evolution. By dynamically allocating attention to relevant regions such as the head, muzzle, or feed trough, the model maintains interpretability while adapting to motion sparsity typical of commercial barns. This allows robust recognition of prolonged behaviors like ruminating and feeding \u0026amp; standing, whose visual cues evolve slowly over time. In our 24/7 barn recordings, TimeSformer achieved 85.0% overall accuracy (macro-F1\u0026thinsp;=\u0026thinsp;0.84) and processed 22.6 fps on an RTX A100, confirming real-time feasibility. The SlowFast architecture, while slightly lower in accuracy (82.3%), remained competitive for posture-oriented states such as Standing or Lying. Its dual-pathway design (slow spatial, fast motion) efficiently captures short-duration transitions like drinking or standing changes, with predictable latency and low memory overhead, making it attractive for cost-constrained edge deployments. Compared with recent benchmarks CBVD-5 \u003csup\u003e37\u003c/sup\u003e(687 segments, 107 Holsteins; ~ 78.7% accuracy) and the beef-cattle dataset of Cao et al. \u003csup\u003e10\u003c/sup\u003e (4974 clips; ~ 90% mAP₅₀) our dataset of 4964 annotated clips (expanded to 9600 via augmentation) encompasses seven behavior classes under natural lighting and occlusion, achieving parity with state-of-the-art accuracy despite far noisier conditions. The scale, behavioral granularity, and continuous capture across diurnal cycles make it one of the largest open dairy-behavior video resources, bridging the gap between controlled research corpora and operational barns.\u003c/p\u003e\u003cp\u003eNonetheless, several methodological limitations merit discussion. Identity leakage remains a primary risk if clips from the same cow appear in both training and test sets, apparent performance may be inflated. Future work will adopt leave-cow-out or day-blocked splits to ensure disjoint identities and report corresponding performance changes. Temporal sampling is another constraint 12 frames per 10-s clip (~\u0026thinsp;1.2 fps) may undersample mandibular cycles critical for rumination and drinking, ablations at 16\u0026ndash;32 frames will clarify the accuracy-latency trade-off. Tracking reliability also warrants quantification, our YOLOv11 and ByteTrack pipeline preserves per-cow IDs but has yet to be evaluated on standard multi-object tracking metrics (MOTA, IDF1, ID-switches) using an annotated subset. Environmental variability further affects performance; accuracy drops 8\u0026ndash;12% at night or under severe occlusion, so stratified tables by day/night, view angle, and occlusion level will be added. Because augmented samples were confined to training, evaluation will be redone on unaugmented test data with natural class priors and balanced accuracy reporting to prevent optimistic bias. The behavioral taxonomy single-label composites such as \u0026ldquo;lying \u0026amp; feeding\u0026rdquo; or \u0026ldquo;standing \u0026amp; ruminating\u0026rdquo; simplifies ethograms that are inherently multi-label. These were chosen for practical monitoring in tie-stall barns, yet future work will re-analyze behavior along orthogonal posture and include confusion plots separating posture from oral-activity errors.\u003c/p\u003e\u003cp\u003eThe structured behavioral outputs cow ID, start/end times, durations, and confidence scores form the behavioral layer of the dairy digital twin. These CSV streams can directly feed NRC feed-intake models to infer dry-matter intake and dynamically adjust rations. Integration within Unity 3D enables real-time visualization of individual cows\u0026rsquo; states synchronized with nutritional and environmental modules, establishing an operational link between perception and physiology. Within the twin dashboard, deviations from baseline feeding or rumination cycles act as early-warning indicators of metabolic or welfare issues, supporting proactive interventions. The architecture will be interoperable with commercial herd-management systems such as DairyComp 305 and Lely T4C and will be designed for modular deployment using networked edge devices. Both TimeSformer and SlowFast achieve real-time throughput on mid-range GPUs, ensuring cost-effective scalability. Economic analyses should consider reductions in manual observation labor, improved feed efficiency, and health-related cost savings. Reliability under barn humidity and dust will be maintained through ruggedized enclosures and modular camera design.\u003c/p\u003e\u003cp\u003eLooking ahead, further improvements will emphasize multi-modal fusion, model calibration, and edge optimization. Acoustic and thermal sensors can complement visual data by capturing rumination chewing sounds and heat signatures, improving discrimination under occlusion or darkness. Lightweight transformer variants using pruning, quantization, and adaptive attention weighting will reduce inference cost and enhance interpretability. Federated-learning frameworks can enable cross-farm model updates without sharing raw video, preserving data privacy while addressing regional variability. Quantitative attention-map analyses reporting the proportion of attention mass within anatomical regions of interest will strengthen biological plausibility. Ultimately, integrating behavioral analytics with NRC feed-intake predictions, health metrics, and environmental sensing will create a self-updating digital-twin ecosystem that links perception, prediction, and management. By demonstrating continuous, per-cow behavioral timelines that synchronize physical animals with their virtual counterparts, this framework establishes a scalable foundation for individualized nutrition, early-warning health systems, and climate-smart, welfare-oriented dairy management.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003e5. Data Availability\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe annotated video dataset (4,964 clips) and augmentation recipes used to create the 9,600‑sample training set include representative sample clips are available on request. Due to facility restrictions, full raw CCTV footage is not publicly released; additional data are available from the corresponding author upon reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e6.\u003c/strong\u003e\u003cstrong\u003e\u0026nbsp;Code Availability\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll scripts for preprocessing, YOLOv11 detection, ByteTrack tracking, and SlowFast/TimeSformer training and inference are available at https://github.com/mooanalytica/digital-twin-dairycow under the MIT license, with a frozen environment file and instructions to reproduce all reported results.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e7. Author contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eS.R. curated data, implemented models, performed experiments, and wrote the manuscript. SR contributed to methodology and analysis. SN supervised the project, provided resources and critical revisions. All authors approved the final version.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e8. Funding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors sincerely thank the Natural Sciences and Engineering Research Council of Canada (RGPIN 2024-04450), the Net Zero Atlantic Canada Agency (300700018), Mitacs Canada (IT36514), and the Department of New Brunswick Agriculture, Aquaculture and Fisheries (NB2425-0025) for funding this study.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e9. Competing interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare no competing financial or non‑financial interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e10. Ethics statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll animal procedures were approved by the Dalhousie University Animal Care and Use Committee (Protocol #2024-026, approval date 16 May 2024) in accordance with Canadian Council on Animal Care (CCAC) guidelines.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e11. Acknowledgments\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors gratefully acknowledge the staff and animal caretakers of the Ruminant Animal Centre at Dalhousie University for their assistance with data collection.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eYoussef, A. \u003cem\u003eet al.\u003c/em\u003e IUMENTA: A generic framework for animal digital twins within the Open Digital Twin Platform. Preprint at https://doi.org/10.48550/arXiv.2411.10466 (2024).\u003c/li\u003e\n\u003cli\u003eEscrib\u0026agrave;-Gelonch, M. \u003cem\u003eet al.\u003c/em\u003e Digital Twins in Agriculture: Orchestration and Applications. \u003cem\u003eJ Agric Food Chem\u003c/em\u003e \u003cstrong\u003e72\u003c/strong\u003e, 10737\u0026ndash;10752 (2024).\u003c/li\u003e\n\u003cli\u003eZhang, Y. \u003cem\u003eet al.\u003c/em\u003e Multimodal Behavior Recognition for Dairy Cow Digital Twin Construction Under Incomplete Modalities: A Modality Mapping Completion Network Approach. Preprint at https://doi.org/10.2139/ssrn.5097747 (2025).\u003c/li\u003e\n\u003cli\u003eZhang, Y. \u003cem\u003eet al.\u003c/em\u003e Digital twin perception and modeling method for feeding behavior of dairy cows. \u003cem\u003eComputers and Electronics in Agriculture\u003c/em\u003e \u003cstrong\u003e214\u003c/strong\u003e, 108181 (2023).\u003c/li\u003e\n\u003cli\u003eRao, S. \u0026amp; Neethirajan, S. Computational Architectures for Precision Dairy Nutrition Digital Twins: A Technical Review and Implementation Framework. \u003cem\u003eSensors\u003c/em\u003e \u003cstrong\u003e25\u003c/strong\u003e, 4899 (2025).\u003c/li\u003e\n\u003cli\u003eBrown-Brandl, T. M. \u0026amp; Tao, J. ASAS-NANP Symposium: mathematical modeling in animal nutrition: harnessing real-time data and digital twins for precision livestock farming. \u003cem\u003eJournal of Animal Science\u003c/em\u003e \u003cstrong\u003e103\u003c/strong\u003e, skaf138 (2025).\u003c/li\u003e\n\u003cli\u003eGuarnido-Lopez, P., Pi, Y., Tao, J., Mendes, E. D. M. \u0026amp; Tedeschi, L. O. Computer vision algorithms to help decision-making in cattle production. \u003cem\u003eAnim Fron\u003c/em\u003e \u003cstrong\u003e14\u003c/strong\u003e, 11\u0026ndash;22 (2024).\u003c/li\u003e\n\u003cli\u003eKurras, F. \u0026amp; Jakob, M. Smart Dairy Farming\u0026mdash;The Potential of the Automatic Monitoring of Dairy Cows\u0026rsquo; Behaviour Using a 360-Degree Camera. \u003cem\u003eAnimals\u003c/em\u003e \u003cstrong\u003e14\u003c/strong\u003e, 640 (2024).\u003c/li\u003e\n\u003cli\u003eAntognoli, V., Presutti, L., Bovo, M., Torreggiani, D. \u0026amp; Tassinari, P. Computer Vision in Dairy Farm Management: A Literature Review of Current Applications and Future Perspectives. \u003cem\u003eAnimals\u003c/em\u003e \u003cstrong\u003e15\u003c/strong\u003e, 2508 (2025).\u003c/li\u003e\n\u003cli\u003eCao, Z. \u003cem\u003eet al.\u003c/em\u003e Semi-automated annotation for video-based beef cattle behavior recognition. \u003cem\u003eSci Rep\u003c/em\u003e \u003cstrong\u003e15\u003c/strong\u003e, (2025).\u003c/li\u003e\n\u003cli\u003eLi, Z. \u003cem\u003eet al.\u003c/em\u003e Method for Dairy Cow Target Detection and Tracking Based on Lightweight YOLO v11. \u003cem\u003eAnimals\u003c/em\u003e \u003cstrong\u003e15\u003c/strong\u003e, 2439 (2025).\u003c/li\u003e\n\u003cli\u003eBai, Q. \u003cem\u003eet al.\u003c/em\u003e X3DFast model for classifying dairy cow behaviors based on a two-pathway architecture. \u003cem\u003eSci Rep\u003c/em\u003e \u003cstrong\u003e13\u003c/strong\u003e, 20519 (2023).\u003c/li\u003e\n\u003cli\u003eZia, A. \u003cem\u003eet al.\u003c/em\u003e CVB: A Video Dataset of Cattle Visual Behaviors. Preprint at https://doi.org/10.48550/arXiv.2305.16555 (2023).\u003c/li\u003e\n\u003cli\u003eNasirahmadi, A., Edwards, S. A., Matheson, S. M. \u0026amp; Sturm, B. Using automated image analysis in pig behavioural research: Assessment of the influence of enrichment substrate provision on lying behaviour. \u003cem\u003eApplied Animal Behaviour Science\u003c/em\u003e \u003cstrong\u003e196\u003c/strong\u003e, 30\u0026ndash;35 (2017).\u003c/li\u003e\n\u003cli\u003eRedmon, J., Divvala, S., Girshick, R. \u0026amp; Farhadi, A. You Only Look Once: Unified, Real-Time Object Detection. in \u003cem\u003e2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)\u003c/em\u003e 779\u0026ndash;788 (IEEE, Las Vegas, NV, USA, 2016). doi:10.1109/CVPR.2016.91.\u003c/li\u003e\n\u003cli\u003eShorten, C. \u0026amp; Khoshgoftaar, T. M. A survey on Image Data Augmentation for Deep Learning. \u003cem\u003eJ Big Data\u003c/em\u003e \u003cstrong\u003e6\u003c/strong\u003e, 60 (2019).\u003c/li\u003e\n\u003cli\u003eZhang, Y. \u003cem\u003eet al.\u003c/em\u003e ByteTrack: Multi-object Tracking by Associating Every Detection Box. in \u003cem\u003eComputer Vision \u0026ndash; ECCV 2022\u003c/em\u003e (eds Avidan, S., Brostow, G., Ciss\u0026eacute;, M., Farinella, G. M. \u0026amp; Hassner, T.) vol. 13682 1\u0026ndash;21 (Springer Nature Switzerland, Cham, 2022).\u003c/li\u003e\n\u003cli\u003eWu, C., Zhou, Y., Pereia Pess\u0026ocirc;a, M. V., Peng, Q. \u0026amp; Tan, R. Conceptual digital twin modeling based on an integrated five-dimensional framework and TRIZ function model. \u003cem\u003eJournal of Manufacturing Systems\u003c/em\u003e \u003cstrong\u003e58\u003c/strong\u003e, 79\u0026ndash;93 (2021).\u003c/li\u003e\n\u003cli\u003eGrant, R. J. \u0026amp; Albright, J. L. Effect of Animal Grouping on Feeding Behavior and Intake of Dairy Cattle. \u003cem\u003eJournal of Dairy Science\u003c/em\u003e \u003cstrong\u003e84\u003c/strong\u003e, E156\u0026ndash;E163 (2001).\u003c/li\u003e\n\u003cli\u003eYu, R. \u003cem\u003eet al.\u003c/em\u003e Research on Automatic Recognition of Dairy Cow Daily Behaviors Based on Deep Learning. \u003cem\u003eAnimals (Basel)\u003c/em\u003e \u003cstrong\u003e14\u003c/strong\u003e, 458 (2024).\u003c/li\u003e\n\u003cli\u003eBorchers, M. R., Chang, Y. M., Tsai, I. C., Wadsworth, B. A. \u0026amp; Bewley, J. M. A validation of technologies monitoring dairy cow feeding, ruminating, and lying behaviors. \u003cem\u003eJournal of Dairy Science\u003c/em\u003e \u003cstrong\u003e99\u003c/strong\u003e, 7458\u0026ndash;7466 (2016).\u003c/li\u003e\n\u003cli\u003eMcHugh, M. L. Interrater reliability: the kappa statistic. \u003cem\u003eBiochem Med (Zagreb)\u003c/em\u003e \u003cstrong\u003e22\u003c/strong\u003e, 276\u0026ndash;282 (2012).\u003c/li\u003e\n\u003cli\u003eDeVries, T. J., Von Keyserlingk, M. A. G., Weary, D. M. \u0026amp; Beauchemin, K. A. Measuring the Feeding Behavior of Lactating Dairy Cows in Early to Peak Lactation. \u003cem\u003eJournal of Dairy Science\u003c/em\u003e \u003cstrong\u003e86\u003c/strong\u003e, 3354\u0026ndash;3361 (2003).\u003c/li\u003e\n\u003cli\u003eBeauchemin, K. A. Invited review: Current perspectives on eating and rumination activity in dairy cows. \u003cem\u003eJournal of Dairy Science\u003c/em\u003e \u003cstrong\u003e101\u003c/strong\u003e, 4762\u0026ndash;4784 (2018).\u003c/li\u003e\n\u003cli\u003eTucker, C. B., Jensen, M. B., De Passill\u0026eacute;, A. M., H\u0026auml;nninen, L. \u0026amp; Rushen, J. Invited review: Lying time and the welfare of dairy cows. \u003cem\u003eJournal of Dairy Science\u003c/em\u003e \u003cstrong\u003e104\u003c/strong\u003e, 20\u0026ndash;46 (2021).\u003c/li\u003e\n\u003cli\u003eJafarigol, E., Trafalis, T. \u0026amp; Mohammadi, N. A Review of Machine Learning Techniques in Imbalanced Data and Future Trends. Preprint at https://doi.org/10.48550/arXiv.2310.07917 (2025).\u003c/li\u003e\n\u003cli\u003eMiftahushudur, T., Sahin, H. M., Grieve, B. \u0026amp; Yin, H. A Survey of Methods for Addressing Imbalance Data Problems in Agriculture Applications. \u003cem\u003eRemote Sensing\u003c/em\u003e \u003cstrong\u003e17\u003c/strong\u003e, 454 (2025).\u003c/li\u003e\n\u003cli\u003eCarreira, J. \u0026amp; Zisserman, A. Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. Preprint at https://doi.org/10.48550/arXiv.1705.07750 (2018).\u003c/li\u003e\n\u003cli\u003eFeichtenhofer, C., Fan, H., Malik, J. \u0026amp; He, K. SlowFast Networks for Video Recognition. in \u003cem\u003e2019 IEEE/CVF International Conference on Computer Vision (ICCV)\u003c/em\u003e 6201\u0026ndash;6210 (IEEE, Seoul, Korea (South), 2019). doi:10.1109/ICCV.2019.00630.\u003c/li\u003e\n\u003cli\u003eBertasius, G., Wang, H. \u0026amp; Torresani, L. Is Space-Time Attention All You Need for Video Understanding? Preprint at https://doi.org/10.48550/arXiv.2102.05095 (2021).\u003c/li\u003e\n\u003cli\u003eIto, K. ASSESSING COW COMFORT USING LYING BEHAVIOUR AND LAMENESS.\u003c/li\u003e\n\u003cli\u003eCardot, V., Le Roux, Y. \u0026amp; Jurjanz, S. Drinking Behavior of Lactating Dairy Cows and Prediction of Their Water Intake. \u003cem\u003eJournal of Dairy Science\u003c/em\u003e \u003cstrong\u003e91\u003c/strong\u003e, 2257\u0026ndash;2264 (2008).\u003c/li\u003e\n\u003cli\u003eAn, J. \u0026amp; Joe, I. Attention Map-Guided Visual Explanations for Deep Neural Networks. \u003cem\u003eApplied Sciences\u003c/em\u003e \u003cstrong\u003e12\u003c/strong\u003e, 3846 (2022).\u003c/li\u003e\n\u003cli\u003eChefer, H., Gur, S. \u0026amp; Wolf, L. Transformer Interpretability Beyond Attention Visualization. in \u003cem\u003e2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)\u003c/em\u003e 782\u0026ndash;791 (IEEE, Nashville, TN, USA, 2021). doi:10.1109/CVPR46437.2021.00084.\u003c/li\u003e\n\u003cli\u003eLi, Z. \u003cem\u003eet al.\u003c/em\u003e Method for Dairy Cow Target Detection and Tracking Based on Lightweight YOLO v11. \u003cem\u003eAnimals\u003c/em\u003e \u003cstrong\u003e15\u003c/strong\u003e, 2439 (2025).\u003c/li\u003e\n\u003cli\u003eBeauchemin, K. A. Invited review: Current perspectives on eating and rumination activity in dairy cows. \u003cem\u003eJournal of Dairy Science\u003c/em\u003e \u003cstrong\u003e101\u003c/strong\u003e, 4762\u0026ndash;4784 (2018).\u003c/li\u003e\n\u003cli\u003eLi, K., Fan, D., Wu, H. \u0026amp; Zhao, A. CBVD-5 (Cow Behavior Video Dataset). Kaggle (2024).\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":false,"email":"","identity":"npj-veterinary-sciences","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"npj Veterinary Sciences","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"Unsupported Journal","inReviewEnabled":false,"inReviewRevisionsEnabled":false},"keywords":"Digital twin, Cattle behavior detection, Deep learning, Computer vision, Precision livestock farming, Dairy monitoring","lastPublishedDoi":"10.21203/rs.3.rs-8032374/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8032374/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eDigital twins in dairy systems require reliable behavioral inputs. We develop a video‑based framework that detects and tracks individual cows and classifies seven behaviors under commercial barn conditions. From 4,964 annotated clips, expanded to 9,600 through targeted augmentation, we couple YOLOv11 detection with ByteTrack for identity persistence and evaluate SlowFast versus TimeSformer for behavior recognition. TimeSformer achieved 85.0% overall accuracy (macro‑F1 0.84) and real‑time throughput of 22.6 fps on RTX A100 hardware. Attention visualizations concentrated on anatomically relevant regions (head/muzzle for feeding and drinking; torso/limbs for postures), supporting biological interpretability. Structured outputs (cow ID, start-end times, durations, confidence) enable downstream use in nutritional modeling and 3D digital‑twin visualization. The pipeline delivers continuous, per‑animal activity streams suitable for individualized nutrition, predictive health, and automated management, providing a practical behavioral layer for scalable dairy digital twins.\u003c/p\u003e","manuscriptTitle":"Video-Based Cattle Behavior Detection for Digital Twin Development in Precision Dairy Systems","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-11-24 12:28:06","doi":"10.21203/rs.3.rs-8032374/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-01-01T15:57:51+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-12-17T12:30:23+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"330738181978091502876431110245029193360","date":"2025-12-10T12:06:23+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-11-24T09:08:04+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-11-20T03:26:11+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"276411199418741462924228095825140168913","date":"2025-11-19T03:22:00+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"160474838674138356075977700638336797332","date":"2025-11-15T13:38:18+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-11-13T08:05:39+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-11-12T08:10:28+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-11-10T08:48:38+00:00","index":"","fulltext":""},{"type":"submitted","content":"npj Veterinary Sciences","date":"2025-11-04T20:52:59+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":false,"email":"","identity":"npj-veterinary-sciences","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"npj Veterinary Sciences","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"Unsupported Journal","inReviewEnabled":false,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"78eae286-a6e6-4acc-863b-6b13704a9a13","owner":[],"postedDate":"November 24th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":58289957,"name":"Biological sciences/Biological techniques"},{"id":58289958,"name":"Biological sciences/Computational biology and bioinformatics"},{"id":58289959,"name":"Physical sciences/Engineering"},{"id":58289960,"name":"Physical sciences/Mathematics and computing"},{"id":58289961,"name":"Biological sciences/Zoology"}],"tags":[],"updatedAt":"2026-02-02T09:26:05+00:00","versionOfRecord":[],"versionCreatedAt":"2025-11-24 12:28:06","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8032374","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8032374","identity":"rs-8032374","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-30T02:00:01.510937+00:00
License: CC-BY-4.0