From Watching to Wishing: A Multimodal Computational Analysis of How PUGC Video-Danmaku Ecology Shapes Travel Intention | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article From Watching to Wishing: A Multimodal Computational Analysis of How PUGC Video-Danmaku Ecology Shapes Travel Intention Feng Ye, Min Yin, Shouqian Sun, Xuanzheng Wang This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8141450/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 9 You are reading this latest preprint version Abstract Digital tourism marketing increasingly relies on platform user-generated content (PUGC), yet mechanisms through which multimodal videos and real-time social interactions shape travel decisions remain largely opaque. This study investigates this "black box" by analyzing 2,650 tourism videos from Bilibili and their synchronized Danmaku (bullet-screen comments) through an innovative computational framework integrating computer vision, natural language processing, and machine learning. Three significant findings emerge: First, cognitive features demonstrate exceptional predictive power for travel intention (42.2% of model importance from only 10.4% of features), while traditional audiovisual features show limited contribution (16.1% importance from 45.8% of features)—questioning visual-centric marketing assumptions. Second, non-linear models achieve substantially higher explanatory power than linear approaches (R² = 0.518 vs. 0.033), revealing threshold effects and cognitive gating mechanisms not captured by additive frameworks. Third, Danmaku appears to function as cognitive scaffolding rather than distraction, with Granger causality suggesting orchestrated attention cascades from visual to audio to textual processing (p < 0.001). By operationalizing previously unobservable mechanisms through extraction of 100 + multimodal features aligned with 475 million Danmaku comments, this study provides new insights into platform affordances and tourism decision-making. The findings suggest important implications for digital tourism strategies: consider prioritizing cognitive activation alongside production aesthetics, facilitating synchronized social interaction, and recognizing that mental engagement may constitute a critical resource in attention economies. PUGC Danmaku travel intention multimodal analysis cognitive load social interaction Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 7 Figure 8 Figure 8 Figure 9 Figure 9 Figure 10 Figure 10 Figure 11 Figure 11 Figure 12 Figure 12 Figure 13 Figure 15 Figure 16 Figure 17 Figure 18 Figure 19 1. Introduction The contemporary tourism landscape is undergoing a fundamental transformation driven by the global shift toward an experience economy where consumers increasingly prioritize authentic, personalized, and transformative travel experiences over standardized tourism products(B. Joseph Pine II and Gilmore James H. 1998; Chen 2025). This evolution has profoundly disrupted traditional marketing approaches, rendering top-down promotional strategies progressively ineffective due to their perceived lack of authenticity and unidirectional communication structure(Fotis et al. 2012; Leung et al. 2013). In response to these shifting dynamics, Professional User-Generated Content (PUGC) has emerged as a dominant force in shaping travel decisions, particularly among younger, digitally native tourists who seek peer validation and authentic experiences over commercial messaging(Pascual-Fraile et al. 2025) PUGC represents a unique hybrid form that strategically combines the production values of professional content with the perceived authenticity and relatability of user-generated materials, creating a powerful influence mechanism that operates through parasocial relationships between creators and viewers(Huang 2024). On platforms like Bilibili and YouTube, these content creators have become critical nodes in the tourism information ecosystem, cultivating trust and influencing millions through carefully crafted multimodal narratives that blur the boundaries between entertainment and marketing(Horton and Richard Wohl 1956; Kim and Ko 2012). Despite the acknowledged importance of PUGC in contemporary tourism marketing, a significant gap persists in our empirical understanding of the specific mechanisms through which this influence operates, with the pathway from passive viewing to concrete travel intention formation remaining largely unexplored(Li and Tu 2024a). This empirical gap becomes particularly pronounced when considering the unique media ecology of platforms like Bilibili, which feature Danmaku, a distinctive form of real-time, synchronous commentary that appears as scrolling text overlays directly on the video screen, transforming solitary viewing into a collective social experience (Yang 2020a; Cauchard et al. 2024). Unlike traditional asynchronous comments that appear below or beside video content, Danmaku creates what scholars term "pseudo-synchronous co-viewing," where viewers experience the illusion of watching together with thousands of others, their reactions and interpretations becoming part of the content itself (Wu et al. 2019; Liu 2024). This technological affordance fundamentally alters the information processing environment in which travel decisions form, introducing continuous social validation signals that may amplify, moderate, or potentially undermine the primary content's influence(Yang et al. 2022). The role of Danmaku in high-involvement decision-making processes like travel planning remains entirely unexplored in the academic literature, despite strong theoretical reasons to expect different dynamics than those observed in low-stakes entertainment or e-commerce contexts where most Danmaku research has focused (Lyu 2021; Fan et al. 2023). Travel decisions involve substantial financial commitments, extended time horizons, and significant personal risks, suggesting that the casual social proof mechanisms effective for impulse purchases may operate differently when viewers contemplate international travel or significant tourism investments. Furthermore, the multimodal complexity of tourism videos, which combines scenic visuals, ambient sounds, narrative voiceovers, and cultural content, creates a rich but potentially overwhelming information environment where Danmaku might serve as either helpful interpretation guides or distracting noise that impedes decision-making. Addressing these critical gaps requires methodological innovation beyond traditional approaches. Survey-based studies and manual content analysis, while valuable for capturing subjective experiences and semantic meanings, cannot adequately capture the dynamic, multimodal, and temporally synchronized nature of PUGC-Danmaku ecology. Recent advances in computational methods, particularly in computer vision, natural language processing, and multimodal machine learning, offer unprecedented opportunities to systematically investigate these complex phenomena at scale while maintaining analytical rigor(Baltrusaitis et al. 2019; Xu et al. 2023). By extracting and analyzing hundreds of features across visual, auditory, and textual modalities, we can move beyond speculation about influence mechanisms to empirical measurement of how specific multimodal configurations and interaction patterns shape viewer responses. This study therefore aims to open the "black box" of PUGC influence by implementing an innovative multimodal computational framework that systematically investigates the complex interplay between video content and real-time audience interaction in shaping travel intentions. Through the analysis of 2,650 tourism videos from Bilibili containing millions of Danmaku comments, we seek to answer three fundamental research questions that address both theoretical and practical concerns in digital tourism marketing. First, we investigate how the multimodal profiles of different types of PUGC travel videos systematically differ, moving beyond intuitive categorizations to data-driven discovery of content strategies. Second, we examine how these distinct multimodal profiles influence viewers' travel intentions, testing whether traditional linear frameworks adequately capture these relationships or whether more complex mechanisms operate. Third, we assess the specific role and relative importance of the Danmaku interaction ecology in this influence process, determining whether synchronized social commentary represents mere digital noise or a critical influence pathway that tourism marketers must understand and leverage. 2. Literature Review 2.1 The Evolution of Tourism Influence in Digital Contexts The transformation of tourism marketing from traditional broadcast approaches to participatory digital narratives represents more than a simple channel shift; it fundamentally alters the mechanisms through which destinations build awareness, shape perceptions, and ultimately influence travel decisions(Pop et al. 2022). Professional User-Generated Content has emerged at the intersection of this transformation, combining professional production capabilities with the authenticity markers (Chen 1986; Luo 1986) that contemporary consumers use to assess credibility in an environment saturated with commercial messaging. While tourism scholars have extensively documented the rise of influencer marketing and its impact on destination image formation (Gretzel 2018), critical gaps remain in our understanding of how PUGC's distinctive affordances, particularly its multimodal orchestration and interactive overlays, mechanistically translate viewing experiences into concrete travel intentions. Recent empirical investigations provide compelling but incomplete evidence for PUGC's effectiveness in tourism contexts. Li (Li and Tu 2024a)and Tu (2024) offer the most direct comparative evidence through experimental manipulation, demonstrating that PUGC's advantage over both amateur user-generated content and professional advertising depends critically on destination type and viewer characteristics. For hedonic destinations emphasizing pleasure and experience, authenticity signals emerged as the dominant influence pathway, while utilitarian destinations focusing on practical benefits gained more from professional production quality. This contingency challenges universal claims about influencer superiority and suggests that the frequently cited "authenticity-professionalism blend" (Peinado and Shim 2024) operates through more nuanced mechanisms than previously theorized. Their finding that credibility and usefulness mediate these effects (Li and Tu 2024b) aligns with dual-processing frameworks from consumer psychology(Samson and Voyer 2012; Wyer and Kardes 2020), yet leaves unexamined the specific audiovisual techniques through which creators achieve this delicate balance between professional polish and authentic expression. The psychological mechanisms underlying PUGC influence extend beyond surface-level credibility assessments to encompass parasocial relationships: the one-sided emotional bonds that viewers develop with media personalities through repeated exposure and perceived intimacy(Horton and Richard Wohl 1956). Contemporary research confirms that these parasocial bonds systematically condition how viewers process both informational and affective content, creating a relationship context that transforms objective destination information into personally relevant travel inspiration (Silaban et al. 2023; Wu et al. 2024). Notably, Nguyen et al. (2024)(Nguyen et al. 2024) provide particularly relevant evidence by demonstrating that inspiration fully mediates the relationship between professionalism and travel planning while only partially mediating sincerity effects, suggesting heterogeneous pathways from creator characteristics to behavioral outcomes. This differentiation becomes crucial for understanding PUGC's influence in high-involvement decisions like travel, where emotional resonance may override rational evaluation of destination attributes(Pham 2007), particularly when viewers have developed strong parasocial bonds with creators(Gong and Li 2017; Liu et al. 2024). However, a critical limitation pervades the existing PUGC tourism literature (Chemin et al. 2025) that undermines theoretical development and practical application. Despite widespread recognition that video represents PUGC's primary medium (Polat et al. 2023) and theoretical claims that multimodal storytelling drives its effectiveness(Shah et al. 2018), empirical studies consistently treat "video" as an undifferentiated, monolithic format without examining the specific affordances that distinguish successful from unsuccessful content. Researchers typically measure exposure to "travel vlogs" or "destination videos" (Buckingham 2009; Ross et al. 2011)without operationalizing the production choices comprising the viewing experience: camera angles, editing pace, music selection, color grading, and narrative structure. This oversight is problematic as production quality shapes viewer perceptions in complex ways. The relationship between amateur and polished presentation styles remains underexplored, with (Ward et al. 2020) Ward et al. (2020) treating production style as a binary category rather than decomposing it into measurable multimodal components. This approach obscures which specific audiovisual elements enhance or undermine authenticity perceptions, limiting our understanding of how production choices influence tourism marketing effectiveness. 2.2 Danmaku as a Unique Mechanism of Synchronous Social Influence The phenomenon of Danmaku introduces an entirely new dimension to digital content consumption that fundamentally challenges traditional models of media influence developed for passive, individual viewing contexts. Originating from Japanese "Niconico" video culture and adapted with distinctive characteristics in Chinese platforms like Bilibili, Danmaku represents more than a technological feature; it constitutes a new form of collective viewing that transforms how audiences process and interpret media content (Fang et al. 2018). Unlike conventional commenting systems where feedback appears spatially separated from content and temporally divorced from specific moments, Danmaku overlays create what Wu et al. (2019)(Wu et al. 2019) term "pseudo-synchronous co-viewing," where comments from different times appear simultaneously on screen, creating the psychological experience of watching alongside a crowd despite physical and temporal separation. The theoretical foundations for understanding Danmaku's influence on viewer cognition and behavior derive from two complementary psychological mechanisms that operate simultaneously during viewing experiences. Emotional contagion theory, first formalized by Hatfield et al. (1993)(Hatfield et al. 1993), predicts that exposure to others' emotional expressions triggers automatic mimicry and affective convergence, leading individuals to "catch" the emotions displayed by those around them. Controlled experimental studies confirm this prediction in Danmaku contexts, with Kamei et al. (2012)(KAMEI et al. 2012) demonstrating measurable affect transfer when viewers are exposed to emotionally congruent overlays, while Zhou et al. (2018)(Zhou et al. 2018) show that excitement-laden Danmaku significantly increases viewers' willingness to send virtual gifts to content creators. These findings suggest that the emotional tone of Danmaku may shape viewers' affective responses to tourism content independent of the primary video's emotional valence, potentially amplifying positive responses to destinations or neutralizing negative impressions through collective enthusiasm. Complementing emotional contagion, social proof theory provides a cognitive pathway through which Danmaku influences decision-making by serving as heuristic cues about content value and appropriate responses(Cialdini 1993). When viewers observe numerous comments expressing desire to visit a destination or praising specific attractions, these expressions function as social validation that reduces uncertainty and legitimizes similar responses. Yuan et al. (2024)(Yuan et al. 2024) provide sophisticated time-varying evidence for this mechanism, showing that Danmaku density bursts correlate with increased engagement metrics even after controlling for content quality, suggesting that the mere presence of active commentary signals content worth attending to. Pan's (2023) (Pan 2023)qualitative analysis reveals that users explicitly interpret Danmaku as "being with" rather than merely observing others, transforming their viewing from isolated consumption into participatory experience where collective reactions guide individual interpretation. However, emerging evidence suggests important boundary conditions that complicate straightforward application of these theories to tourism contexts((Brian) Lin et al. 2022). The vast majority of Danmaku research has examined low-stakes, hedonic consumption environments, such as gaming streams, variety shows, and e-commerce broadcasts (Chen et al. 2015; Pan 2023), where decisions involve minimal risk and immediate gratification. In these contexts, the "lively atmosphere" created by dense Danmaku reliably increases engagement, purchase intention, and platform loyalty (Fang et al. 2018; Yuan et al. 2024). Yet Zhang and Ruan (2024)(Rahwan et al. 2024) recently uncovered a critical counter-effect that challenges universal application of social proof principles: excessive comment homogeneity breeds psychological reactance, where viewers resist perceived social pressure and assert autonomy by rejecting popular opinions. This reactance effect may be particularly pronounced in high-involvement decisions like travel planning, where substantial financial and temporal investments make viewers more sensitive to manipulation attempts and less susceptible to momentary social influence. 2.3 Methodological Limitations and the Computational Turn The complexity inherent in PUGC-Danmaku ecology exposes fundamental limitations in tourism research's traditional methodological toolkit, revealing a troubling disconnect between theoretical sophistication and empirical capability. Contemporary theories emphasize multimodal orchestration, temporal dynamics, and real-time social interaction as defining features of digital tourism influence, yet the field's dominant methods, namely surveys, interviews, and manual content analysis, were designed for static, monomodal data and cannot adequately capture these phenomena (Naab et al. 2019). This methodological constraint has created a literature rich in conceptual frameworks but impoverished in mechanistic evidence, where researchers theorize about complex influence processes they cannot empirically observe or measure. Survey-based studies, while remaining the dominant approach in tourism research, suffer from systematic validity threats that become particularly acute when investigating digital media consumption(Luo et al. 2021; Otto et al. 2022). Naab et al. (2019) provide sobering evidence that retrospective self-reports of even simple behaviors like video viewing duration show average errors of 40–60% when compared to server logs, with systematic biases toward overestimation of socially desirable behaviors and underestimation of passive consumption. For complex multimodal experiences involving rapid scene changes, background music, narrative voiceovers, and scrolling text overlays, the cognitive burden of accurate recall likely increases exponentially, rendering self-report data highly suspect(Pink and Newton 2020). More fundamentally, surveys cannot capture the temporal dynamics central to digital influence: the specific moments when visual revelations align with musical crescendos and collective commentary to create peak emotional experiences that crystallize travel intentions. Manual content analysis faces even more severe limitations when applied to PUGC-Danmaku ecology, as the sheer volume and complexity of data overwhelm human coding capacity(Grimmer and Stewart 2013; Alaei et al. 2019). While sophisticated frameworks like Schwenzow et al.'s (2021)(Schwenzow et al. 2021) multimodal coding scheme provide systematic protocols for analyzing video content across multiple dimensions, the authors acknowledge that labor requirements increase "immensely" with video length and modal complexity(Peinado et al. 2020), rendering them impractical for the thousands of hours typical in PUGC corpora. The challenge compounds when considering Danmaku, where a single popular video may contain hundreds of thousands of time-coded comments that would require years of manual coding to analyze comprehensively. Attempts to address scale through crowdsourcing sacrifice the interpretive nuance and contextual understanding that justify human coding over automated methods(Larson et al. 2014), while inter-coder reliability deteriorates rapidly as the number of variables and modalities increases(Bayerl and Paul 2011). The emergence of computational methods, particularly advances in deep learning and multimodal machine learning, offers a principled solution to these methodological constraints while enabling investigation of previously unobservable phenomena. Contemporary computer vision models can extract hundreds of visual features from video frames, including object detection, scene classification, aesthetic attributes, and emotional expressions, with superhuman accuracy and perfect reliability (He et al. 2016). Natural language processing techniques can analyze millions of comments to identify semantic patterns, emotional valences, and social network structures that would be impossible to detect manually(Vaswani et al. 2023). Most critically, multimodal fusion architectures can model the complex interactions between visual, auditory, and textual streams, capturing cross-modal dependencies and temporal dynamics that linear statistical methods cannot represent(Baltrusaitis et al. 2019; Xu et al. 2023) . 2.4 An Integrated Theoretical Framework: The Computationally Extended S-O-R Model Building on the foundational Stimulus-Organism-Response paradigm that has guided environmental psychology and consumer behavior research for decades (Mehrabian and Russell 1980), we propose a Computationally Extended S-O-R Model specifically designed to capture the unique dynamics of PUGC-Danmaku influence in tourism contexts. This framework addresses critical limitations in existing S-O-R applications to digital tourism, which typically treat stimuli as static and independent(Asyraff et al. 2023), assume linear additive effects, and rely on self-reported organism states that suffer from severe measurement error(Li et al. 2015; Yüksel 2017). Our extension integrates insights from cognitive load theory(Sweller 1988), social proof theory, and parasocial relationship theory(Horton and Richard Wohl 1956) while enabling empirical validation through computational measurement of previously unobservable constructs. The stimulus layer in our framework recognizes that viewers encounter not a single coherent message but multiple competing information streams that vie for limited cognitive resources(Turner and O’Leary 2012). Primary video content delivers multimodal stimuli through visual semantics captured by CLIP embeddings(Radford et al. 2003), acoustic features including spectral characteristics and temporal patterns, and narrative structures reflected in speech transcription and pacing. Simultaneously, Danmaku overlays provide continuous social stimuli through comment density patterns, semantic content, emotional valence, and temporal clustering that may reinforce, contradict, or reframe the primary content's message. Unlike traditional S-O-R applications that assume stimuli combine additively(Thein et al. 2022), our framework explicitly models modal competition through cross-modal correlation analysis and temporal alignment metrics, recognizing that simultaneous information streams may interfere with rather than enhance each other. The organism layer represents the critical mediating processes through which external stimuli translate into behavioral responses, but rather than relying on error-prone self-reports(Podsakoff and Organ 1986; Stone and Shiffman 2002), we infer cognitive and affective states from objective behavioral traces. Cognitive load, traditionally measured through subjective ratings(Skulmowski and Rey 2015) or secondary task performance(Haji et al. 2015; Ehlers 2020), emerges from the computational complexity of multimodal features(Brunken et al. 2003; Ross et al. 2003), wherein the variance in visual semantics indicates intrinsic load from content difficulty, cross-modal asynchrony reflects extraneous load from poor design(Kan 2023), and the semantic coherence between video and Danmaku suggests germane load from meaningful elaboration(Leng et al. 2016). Social processing manifests through Danmaku clustering patterns that reveal when viewers collectively attend to specific moments, sentiment cascades that show emotional contagion(Yu et al. 2025), and semantic convergence that indicates shared interpretation. Parasocial bonding, while not directly observable, leaves traces in the personalization of comments, frequency of creator references, and persistence of viewing across a creator's catalog(Fazli-Salehi et al. 2022; Tan-intaraarj 2024). The response layer captures behavioral outcomes through platform-native indicators rather than artificial research instruments, enhancing ecological validity(SU et al. 2021) while enabling large-scale measurement(Diehl et al. 2017). Travel intention, our primary outcome, emerges from computational analysis of viewer comments(Mehra 2023) where expressions like "I want to visit" and "added to my bucket list" represent natural behavioral signals uncontaminated by research demand effects(Caulley 1994; Bailenson et al. 2004). Engagement patterns including viewing duration, replay behavior, and sharing actions provide complementary indicators of content influence(Wu et al. 2018; Anh 2024), while the temporal distribution of intention expressions reveals which specific moments trigger decision crystallization(Li et al. 2025). This behavioral approach acknowledges that expressed interest may not equal booking behavior but captures the crucial early stages of travel decision-making where awareness transforms into consideration(Dimitriou and AbouElgheit 2019). 2.5 Research Hypotheses Development The theoretical foundations and methodological innovations discussed above converge on fundamental questions about how PUGC videos and their Danmaku overlays jointly shape travel intentions through multimodal orchestration and social dynamics. Our hypothesis development proceeds from three interconnected theoretical gaps that emerged from the literature review. First, despite recognition that PUGC creators employ different communication strategies(Pertiwi and Sanusi 2023), no empirical evidence exists for whether these strategies manifest as measurably distinct multimodal profiles or represent random variation(Zhu et al. 2025). Second, while theory predicts that multimodal features and cognitive processing should influence travel decisions(Jun and Vogt 2013), the functional form of these relationships, whether linear, threshold-based, or interactive, remains unknown. Third, although Danmaku theoretically provides social influence, its relative importance compared to primary content features in high-involvement decisions like travel has never been tested(Filieri et al. 2025). Building on cognitive load theory's prediction that information complexity systematically affects processing and retention (Sweller 1988), combined with evidence that different tourism content serves distinct communication goals, we propose that PUGC videos will cluster into categories with characteristic multimodal signatures. These categories should differ not only in surface features like topic and style but in fundamental properties including visual dynamism, audio characteristics, narrative structure, and crucially, the cognitive demands they impose on viewers. Furthermore, these differences should manifest in how audiences engage through Danmaku, with some content types triggering immediate reactive commentary while others promote reflective discussion. Therefore: H1: Data-driven categories of PUGC travel videos will exhibit systematically different multimodal and cognitive profiles that reflect distinct communication strategies. H1a: Categories will show significant differences in objective multimodal features including visual dynamism, audio characteristics, and Danmaku density patterns, demonstrating that content types employ distinct production approaches rather than random variation. H1b: Categories will impose different levels of cognitive load as computationally derived from multimodal feature complexity and temporal variation, indicating that cognitive demands represent a strategic choice rather than incidental byproduct. The relationship between multimodal features and travel intention likely operates through complex mechanisms that traditional linear models cannot capture. Cognitive load theory suggests threshold effects where information must reach sufficient complexity to engage deep processing but not so much as to overwhelm capacity. Social proof theory predicts that Danmaku influence depends on reaching critical mass where collective enthusiasm becomes self-reinforcing. Parasocial relationship theory implies that influence accumulates through repeated exposure to consistent creator personas rather than linearly with each video. These theoretical predictions, combined with evidence of non-linear effects in related domains, suggest: H2: The multimodal and cognitive profiles of PUGC videos will significantly predict viewers' travel intentions through identifiable causal pathways. H2a: Multimodal feature dimensions and cognitive load metrics will demonstrate significant associations with travel intention expressions, though these relationships may be non-linear and involve threshold effects invisible to traditional regression approaches. H2b: Temporal causality analysis will reveal systematic lead-lag relationships between content features and audience reactions, with visual elements triggering audio responses that subsequently generate Danmaku commentary, providing evidence for orchestrated influence sequences. The integration of multiple information streams in PUGC-Danmaku ecology suggests that influence emerges from the gestalt rather than individual components. Multimodal learning theory demonstrates that cross-modal integration often yields superior outcomes compared to single modalities when information streams are complementary rather than redundant. In tourism contexts, visual beauty might capture attention, narrative provides meaning, music sets emotional tone, and Danmaku offers social validation, with each modality being insufficient alone but powerful in combination. Moreover, the unique cognitive demands of processing simultaneous video and text streams may create distinctive influence patterns where cognitive engagement becomes the bottleneck determining whether multimodal richness translates into behavioral influence. Therefore: H3: Multimodal integration will provide superior prediction of travel intention compared to unimodal approaches, with interactive features and cognitive processing playing critical roles. H3a: Machine learning models using integrated multimodal features will significantly outperform single-modality baselines in predicting travel intention expressions, demonstrating that influence emerges from cross-modal interactions rather than additive effects. H3b: Feature importance analysis will reveal that audience interaction metrics and cognitive load measures are among the most influential predictors, potentially surpassing traditional content quality indicators like visual aesthetics or production values. 3. Methodology 3.1 Data Collection and Computational Framework This investigation employed a comprehensive computational framework to analyze PUGC tourism videos and their associated Danmaku commentary from Bilibili, China's leading video-sharing platform characterized by its distinctive synchronous commenting system. The selection of Bilibili as our data source was motivated by three factors that make it ideal for investigating multimodal influence mechanisms: the platform's native integration of Danmaku creates a natural laboratory for studying synchronized social viewing, its predominantly young user base (82% aged 18–35) represents the digitally native tourists driving industry transformation, and its open API enables systematic data collection at scales impossible with manual methods. The data collection proceeded through multiple phases between January and February 2025, beginning with exploratory sampling to understand content diversity and concluding with targeted collection to ensure adequate representation across emergent categories. This study employed a comprehensive computational framework to analyze PUGC tourism videos from Bilibili, China's leading video-sharing platform characterized by its distinctive Danmaku commenting system. The platform's unique ecology, where time-synchronized comments overlay video content, provides an ideal natural laboratory for investigating how multimodal content and real-time social interactions jointly influence travel decisions. Data collection proceeded through two complementary phases between January and February 2025. The initial exploratory phase cast a wide net using broad tourism-related Chinese keywords ("旅游" [tourism], "旅行" [travel], "游玩" [leisure travel]) to understand the landscape of tourism content, yielding metadata for 1,640 videos. After quality screening removed duplicates and videos shorter than 60 seconds, 1,400 videos remained for analysis. To avoid imposing researcher-defined categories while ensuring systematic organization, we employed unsupervised clustering via BERTopic(Grootendorst 2020). Video titles and descriptions underwent segmentation using jieba with a custom tourism vocabulary, then were embedded into 1,024-dimensional semantic space using Conan-embedding-v1. The clustering pipeline utilized UMAP for dimensionality reduction (n_neighbors = 15, n_components = 10, min_dist = 0.1) to preserve both local and global structure, followed by HDBSCAN (min_cluster_size = 30, min_samples = 10) for density-based cluster identification. This process revealed 67 fine-grained topics that were hierarchically consolidated into three macro-categories using Ward's linkage, representing distinct communication strategies in PUGC tourism content. The emergent categorization revealed imbalanced representation across categories, motivating a targeted second collection phase using category-specific keywords. This yielded a final corpus of 2,650 videos: 991 Destination Attraction videos focusing on scenic locations, 1,116 Tourism Narrative Communication videos emphasizing personal experiences, and 543 Overseas Cultural Experience videos highlighting international encounters. For each video, we collected complete video files, all Danmaku comments with precise timestamps (N = 475,253,138), and hierarchical viewer comments (N = 11,849,985), creating a rich multimodal dataset for analysis. 3.2 Multimodal Feature Extraction and Engineering Feature extraction required balancing temporal granularity with computational efficiency while preserving content dynamics. We implemented a sliding window approach with 5-second windows and 2.5-second overlap, a configuration that captures content transitions while maintaining sufficient stability for reliable feature extraction. This approach generated overlapping segments that preserve temporal continuity, essential for understanding how multimodal patterns evolve throughout videos. Visual feature extraction leveraged three complementary approaches to capture semantic and aesthetic properties. CLIP ViT-B/32 generated 512-dimensional semantic embeddings for each frame, enabling quantification of visual diversity and thematic coherence. ResNet50 pre-trained on Places365 identified environmental contexts from 365 scene categories particularly relevant to tourism. YOLOv8n provided object detection for measuring visual complexity through entity counts. Aggregation from frame-level to window-level employed both statistical summarization (mean, standard deviation, range) and temporal modeling (trend coefficients, change rates), yielding 627 visual feature dimensions that comprehensively characterize visual content. Audio analysis extracted 75 acoustic features via librosa, encompassing multiple perceptual dimensions. These included 13 MFCCs with first and second derivatives for timbral characterization essential for speech-music discrimination, spectral features (centroid, bandwidth, contrast, rolloff, flatness) differentiating ambient soundscapes from narration, temporal features (zero-crossing rate, tempo, onset patterns) indicating activity levels, and tonal features (chroma vectors, harmonic-percussive separation) distinguishing musical accompaniment from environmental sounds. This comprehensive acoustic profile captures production quality, emotional tone, and information density independent of semantic content. Danmaku processing addressed the unique challenges of time-synchronized, overlapping text streams. Within each 5-second window, we aggregated all comments while preserving their temporal context. Text preprocessing employed jieba segmentation optimized for informal online language, supplemented by custom dictionaries for tourism terminology and internet slang. Each window's aggregated Danmaku was characterized through 1,024-dimensional embeddings (Conan-embedding-v1), density metrics (raw count, unique users, temporal clustering coefficients), and linguistic features (sentiment scores, lexical diversity, semantic coherence with video content). These features capture both quantitative and qualitative aspects of collective viewing experiences. 3.3 Multimodal Fusion and Cognitive Load Operationalization The integration of heterogeneous modalities required sophisticated fusion techniques that preserve modality-specific information while capturing cross-modal interactions. Our two-stage fusion architecture addresses the fundamental challenge of comparing features from different measurement spaces while modeling their complex interdependencies. Stage one employed Deep Canonical Correlation Analysis (DCCA) to learn aligned representations between visual and audio modalities. The architecture consisted of parallel neural networks processing visual features (627 dimensions) and audio features (75 dimensions), both utilizing ReLU activation and dropout regularization. Training optimized these networks to maximize correlation between their 256-dimensional output representations. Stage two integrated the aligned audiovisual representations with Danmaku embeddings using a Transformer architecture that captures complex inter-modal dependencies through self-attention mechanisms. With eight attention heads enabling simultaneous focus on different relationship aspects and three encoder layers providing hierarchical representation learning, this architecture produced 256-dimensional fused representations. Layer normalization and dropout (p = 0.1) ensured stable training and generalization, yielding holistic multimodal representations for each temporal window. Building on Cognitive Load Theory's three-component framework (Sweller 1988), we operationalized cognitive dimensions through computational metrics. Intrinsic load, representing content's inherent difficulty, was measured through three components: Content Complexity (CC) : $$\:\begin{array}{c}CC=\frac{\sigma\:\left({F}_{f}\right)}{\mu\:\left({F}_{f}\right)}\end{array}$$ 1 Semantic Density (SD) : $$\:\begin{array}{c}SD=\text{Var}\left({F}_{t}\right)\end{array}$$ 2 Conceptual Difficulty (CD) : $$\:\begin{array}{c}CD=\frac{\text{l}\text{o}\text{g}(DI\times\:DD+1)}{10}\end{array}$$ 3 where F f represents fused features, F t represents text-aligned features, σ denotes standard deviation, µ denotes mean, Var denotes variance, DI is demand intensity, and DD is demand diversity. Extraneous load, representing presentation-imposed difficulty, was captured through: Modal Interference (MI) : $$\:\begin{array}{c}MI=1-\left|\rho\:\right({F}_{v},{F}_{a}\left)\right|\end{array}$$ 4 Presentation Complexity (PC) : $$\:\begin{array}{c}PC=\sigma\:\left({F}_{v}\right)\end{array}$$ 5 where F v represents video-aligned features, F a represents audio-aligned features, and ρ denotes Pearson correlation coefficient. These operationalizations transform theoretical constructs into measurable quantities suitable for empirical analysis. 3.4 Travel Intention Measurement Through Computational Text Analysis Measuring travel intention in natural viewing contexts required innovation beyond traditional survey approaches. We developed a computational method to extract intention signals from 11.8 million viewer comments, providing unprecedented scale while maintaining ecological validity since comments represent spontaneous expressions uninfluenced by research instruments. Our dual-stage approach balanced coverage with precision. The rule-based stage employed 47 carefully developed patterns capturing Chinese travel intention expressions, including direct statements ("想去" [want to go], "种草了" [added to list]), planning language ("打算去" [planning to visit], "安排上" [scheduling it]), and destination inquiries ("在哪里" [where is this], "怎么去" [how to get there]). However, keyword matching alone produced substantial false positives from conditional statements, quotations, and sarcasm. Therefore, a second-stage RoBERTa model fine-tuned on manually annotated comments distinguished genuine intentions from other travel-related language use. The final intention score was calculated as: Travel Intention Score (TIS) : $$\:\begin{array}{c}TI{S}_{i}=\frac{{V}_{i}}{{C}_{i}}\end{array}$$ 6 where V i represents validated intention expressions for video i and C i represents total comments for video i . This provides a continuous measure of the proportion of viewers moved to express travel interest. 3.5 Statistical Analysis and Model Development Prior to analysis, we conducted comprehensive assumption diagnostics. Homogeneity of variances was assessed using Levene's test for subsequent ANOVA procedures. Regression diagnostics included Shapiro-Wilk tests for residual normality, residual versus fitted value plots for homoscedasticity evaluation, and Variance Inflation Factors for multicollinearity detection. When violations were detected, we employed appropriate robust methods or non-parametric alternatives. To examine differences in multimodal profiles across video categories, one-way ANOVA was conducted for each of the 102 extracted features, with video category as the independent variable. Effect sizes were calculated using eta-squared to assess practical significance, and post-hoc comparisons employed Tukey's HSD test. The Benjamini-Hochberg procedure controlled false discovery rate at 0.10 to balance Type I error control with statistical power. To investigate the relationship between multimodal features and travel intention, we specified a multiple linear regression model using selected key features: $$\:\begin{array}{c}{Y}_{i}={\beta\:}_{0}+{\beta\:}_{1}{F}_{i1}+{\beta\:}_{2}{F}_{i2}+{\beta\:}_{3}{C}_{i1}+{\beta\:}_{4}{C}_{i2}+{\delta\:}_{1}{D}_{i1}+{\delta\:}_{2}{D}_{i2}+{\epsilon\:}_{i}\end{array}$$ 7 where Y i represents travel intention for video i , F i1 and F i2 denote fused feature statistics (mean activation and temporal consistency respectively), C i1 and C i2 represent cognitive load measures (intrinsic content complexity and extraneous modal interference respectively), D i1 and D i2 are dummy variables for video categories (with Destination Attraction as reference category), and ε i is the error term. Given identified assumption violations, robust standard errors were computed to ensure valid inference. Temporal causality among modalities was examined using Granger causality tests with a maximum lag of five windows, preceded by Augmented Dickey-Fuller tests to verify stationarity. Recognizing that digital influence mechanisms may operate through non-linear pathways, we implemented Random Forest regression to capture complex relationships and interactions. To prevent overfitting and ensure generalizability, feature selection was performed exclusively on the training set (75% of data), removing features with correlations exceeding 0.98 to eliminate redundancy. Model hyperparameters were conservatively configured (100 estimators, max_depth = 8, min_samples_split = 10, max_features='sqrt') to balance complexity with generalization. Performance evaluation employed 5-fold cross-validation on the training set and held-out test set validation, with bootstrap resampling (1,000 iterations) providing confidence intervals for performance metrics. Model interpretation utilized SHAP (SHapley Additive exPlanations) analysis to decompose predictions into feature contributions, revealing both global importance patterns and local decision boundaries. All analyses were implemented in Python 3.9 using scikit-learn (1.3.0) for machine learning, statsmodels (0.14.0) for statistical testing, PyTorch (2.0.1) for deep learning components, and SHAP (0.42.1) for model interpretation. Computations were performed on hardware featuring NVIDIA RTX 3090 GPU and 64GB RAM to ensure computational efficiency. 4. Results 4.1 Data-Driven Categorization of Tourism Video Content The computational analysis of 2,650 tourism-themed videos from Bilibili through BERTopic modeling revealed a hierarchical structure of 67 distinct topics, subsequently consolidated into three primary categories via agglomerative clustering with Ward's linkage method. The hierarchical clustering dendrogram (Fig. 1 ) demonstrated clear semantic boundaries between topic clusters, enabling robust categorization into: Destination Attraction videos (37.4%, n = 991), Tourism Narrative Communication content (42.1%, n = 1,116), and Overseas Cultural Experience materials (20.5%, n = 543) (Table 1 ). Table 1 Distribution and Characteristics of Video Categories Category N % Representative Topics Key Lexical Features Destination Attraction 991 37.4 "Qingdao Seaside Travel", "Hokkaido Winter Snow", "Tibet Self-Driving" Geographic names, visual descriptors, spatial terms Tourism Narrative Communication 1,116 42.1 "Leisure Travel Log", "Youth Group Tour", "Travel Guide Alone" Personal pronouns, temporal markers, emotional language Overseas Cultural Experience 543 20.5 "Travel News in India", "Southern Europe Vlog", "Korean Girls Travel to Shanghai" Cultural terms, international place names, comparative language The validity of this categorization was quantitatively confirmed through similarity matrix analysis (Fig. 2 ), which demonstrated a pronounced block-diagonal structure. Within-category similarity scores (M = 0.82, SD = 0.07) significantly exceeded between-category similarities (M = 0.64, SD = 0.11), yielding a substantial effect size (t(2648) = 47.28, p < 0.001, Cohen's d = 1.84), confirming genuine categorical distinctions rather than arbitrary divisions. The intertopic distance map (Fig. 3 ) further corroborated categorical distinctions, with each category occupying distinct regions in the reduced dimensional space. Temporal analysis from 2018 to 2024 (Fig. 4 ) revealed differential growth trajectories: Overseas Cultural Experience demonstrated the most pronounced expansion (CAGR = 35.2%), particularly accelerating post-2023 coinciding with relaxed international travel restrictions, while Destination Attraction showed steady growth (CAGR = 11.2%). 4.2 Multimodal and Cognitive Profiles of Video Categories To test H1, we conducted systematic analyses of variance across 102 computationally extracted features. The analysis revealed that 94 of 102 features (92.2%) demonstrated significant categorical variation at p < 0.05 after Benjamini-Hochberg correction for multiple comparisons (Table 2 ). Table 2 Summary of ANOVA Results Across Feature Domains Feature Domain N Features Significant (%) Mean η² Max η² Audio Features 89 92.1% 0.044 0.072 Visual Features 8 87.5% 0.031 0.052 Danmaku Interaction 2 100% 0.048 0.063 Cognitive Load* 5 - - - Note: While Levene's tests indicated heterogeneous variances for some features (p < 0.05), the large effect sizes and consistent patterns across multiple features support the robustness of our categorical distinctions. Welch's ANOVA confirmed all significant findings. Danmaku interaction patterns revealed striking categorical differences (Fig. 5 ). Destination Attraction videos elicited the highest commentary density (M = 1.50 per window, SD = 5.89), significantly exceeding Overseas Cultural Experience videos (M = 0.73, SD = 4.85; F(2, 2647) = 89.34, p < 0.001, η² = 0.063). This pattern suggests differentiated audience engagement strategies, with scenic content triggering immediate reactions while cross-cultural content promotes contemplative viewing. Comprehensive acoustic profiling across 89 features revealed distinctive sonic fingerprints for each category (Fig. 6 ). Spectral centroid analysis demonstrated that Destination Attraction videos contained the brightest audio (M = 1002.19 Hz), while MFCC analysis showed Overseas Cultural Experience videos had substantially lower spectral energy (MFCC 0: M = -175.48), suggesting distinct production contexts. Visual analysis revealed that Tourism Narrative Communication videos contained significantly more detected objects (M = 32.40, SD = 49.11) compared to other categories (F(2, 2647) = 52.78, p < 0.001, η² = 0.038), while CLIP embedding magnitudes indicated Overseas Cultural Experience videos possessed richer semantic content (Fig. 7 ). 4.3 Limitations of Linear Modeling: An Exploratory Analysis Important Note on Statistical Assumptions Before presenting the linear regression results, we must acknowledge severe violations of statistical assumptions that fundamentally limit their interpretation. Diagnostic tests revealed Extreme non-normality (Jarque-Bera = 7,742,617, p < 0.001; skewness = 11.89, kurtosis = 266.74) Severe multicollinearity with VIF values exceeding 100 for some features (condition number = 487) Influential observations (maximum Cook's distance = 0.2179) These violations render traditional linear inference unreliable. We present these results solely for completeness and as motivation for the non-linear approaches that follow. The linear model's minimal explanatory power (R² = 0.033, Table 3 ) further confirms the inadequacy of linear frameworks for these data. Regression diagnostic plots (Fig. 8 ) visually confirm these severe assumption violations.. Table 3 Multiple Regression Results for Travel Intention Variable β SE t p 95% CI Constant 9.79 44.09 0.22 0.824 [-76.66, 96.24] Fused Temporal Consistency -5.20 0.80 -6.50 < 0.001*** [-6.77, -3.63] Category: Overseas Cultural Experience -4.59 0.91 -5.06 < 0.001*** [-6.37, -2.81] Fused Statistical Mean Activation -46.70 31.85 -1.47 0.143 [-109.15, 15.75] Cognitive Intrinsic Load Content Complexity 13.43 37.55 0.36 0.721 [-60.20, 87.06] Cognitive Extraneous Load Modal Interference 3.07 9.04 0.34 0.734 [-14.65, 20.79] Category: Tourism Narrative Communication 1.13 1.26 0.90 0.367 [-1.33, 3.60] Note: ***p < 0.001; Results should be interpreted as exploratory patterns only due to assumption violations. Figure 8 : Regression Diagnostic Plots Despite these limitations, Granger causality tests revealed potentially interesting temporal patterns. For Overseas Cultural Experience videos (5,430 windows analyzed), visual changes preceded audio responses in 14.3% of windows (p < 0.05), which subsequently triggered Danmaku commentary in 13.5% of windows, creating a visual-to-audio-to-text cascade (χ² = 287.43, p < 0.001). 4.4 Non-linear Modeling and Feature Importance Analysis Given the failure of linear approaches, we implemented Random Forest regression with careful attention to overfitting concerns. Systematic ablation studies across feature configurations revealed a striking pattern: cognitive features dramatically outperformed traditional multimodal features in predicting travel intention (Table 4 ). Table 4 Ablation Study Results Model Configuration N Features Test R² Train R² Overfitting Gap MAE Interpretation Cognitive Features Only 5 0.518 0.948 0.430* 2.10 High risk of overfitting All Multimodal Features 48 0.366 0.822 0.456* 4.74 Severe overfitting Text Features Only 7 0.144 0.525 0.381* 7.71 Moderate overfitting Fused Features Only 12 0.084 0.436 0.352* 8.47 Moderate overfitting Visual Features Only 12 0.078 0.418 0.340* 8.52 Moderate overfitting Audio Features Only 10 0.039 0.382 0.343* 9.47 Moderate overfitting *Note: Overfitting gaps > 0.3 indicate substantial overfitting, requiring cautious interpretation. The model complexity versus performance analysis (Fig. 9 ) revealed while the cognitive features model achieved the highest test R² (0.518), the substantial train-test gap (0.430) suggests the model may be capturing idiosyncratic patterns rather than generalizable relationships. This finding supports H3a (cognitive features dominate prediction) but with important caveats regarding generalizability. Interestingly, the low canonical correlation between audio and visual features (CCA = 0.028) suggests weak inherent alignment in PUGC content, further supporting our finding that cognitive processing rather than multimodal harmony drives travel intention formation. This weak audiovisual correlation contradicts assumptions about professional content quality and indicates that PUGC creators may not optimize cross-modal coherence, making cognitive interpretation by viewers the critical determinant of influence. SHAP analysis revealed that 'cognitive intrinsic load conceptual difficulty' dominated with 40.1% of total importance, followed by 'text temporal variability' (11.9%) and 'fused temporal consistency' (8.4%). The concentration of importance in cognitive features (42.2% of importance from 10.4% of features) suggests that travel intention formation depends primarily on cognitive processing rather than sensory richness, though this interpretation must be tempered by overfitting concerns. SHAP dependence analysis (Fig. 11 ) revealed complex non-linear relationships with clear threshold effects and saturation patterns, explaining why linear models failed. These non-monotonic patterns, while compelling, require validation on independent datasets given the overfitting indicators. The final optimized model's performance (Fig. 12 ) demonstrates the final optimized model, using conservative hyperparameters to reduce overfitting, achieved Test R² = 0.378 (95% Bootstrap CI: [0.249, 0.507]) with MAE = 4.33. While showing lower R² than the pure cognitive model, this configuration demonstrated a more acceptable train-test gap (0.31), suggesting better generalizability. 5. Discussion This investigation reveals a fundamental reconceptualization of digital tourism influence mechanisms within platform-mediated environments. Analysis of 2,650 PUGC videos uncovers three transformative insights that challenge conventional understanding of tourism marketing in digital contexts, as summarized in Table 5 . Cognitive processing demonstrates primacy over sensory stimulation in determining travel intention formation, with cognitive features dominating predictive importance despite comprising a small fraction of analyzed features. Furthermore, multimodal influence operates through non-linear threshold mechanisms rather than additive effects, evidenced by the dramatic performance disparity between linear and non-linear models. Additionally, synchronized social commentary through Danmaku functions as cognitive scaffolding, transforming passive viewing into collective sense-making processes with distinct categorical patterns. Table 5 Summary of Hypothesis Testing Results Hypothesis Result Key Evidence Theoretical Implications H1a Strongly Supported 92.2% of features showed significant variation (p < 0.05) PUGC categories represent distinct multimodal communication strategies H1b Supported Cognitive features showed significant categorical differences Content complexity systematically varies across tourism communication goals H2a Not Supported Linear model R² = 0.033; assumption violations Linear frameworks fundamentally inadequate for PUGC data H2b Supported Granger causality confirmed (χ² = 287.43, p < 0.001) Predictable influence cascades exist across modalities H3a Partially Supported Non-linear R² = 0.378 (optimized) vs. linear R² = 0.033; overfitting concerns with higher R² models Non-linear integration improves prediction but generalizability remains challenging H3b Strongly Supported Cognitive features 42.2% importance with 10.4% of features Cognitive processing and interaction dominate over content features Note: H3a's R² = 0.518 represents the cognitive-only model with substantial overfitting (train-test gap = 0.430). The final optimized model achieved R² = 0.378 with reduced overfitting. 5.1 The Cognitive Primacy Revolution in Digital Tourism Influence Mechanisms The sixfold predictive advantage of cognitive features (R² = 0.518) relative to audiovisual features (R² < 0.08) challenges foundational assumptions underlying tourism marketing since Urry's (1990)(Urry 1990) conceptualization of the tourist gaze, specifically the presumption that visual stimulation drives travel desire. Although observed overfitting (train-test gap = 0.430) necessitates interpretive caution, the pattern's consistency across multiple model specifications suggests genuine phenomena whereby cognitive engagement mediates and potentially supersedes sensory appeal within interactive digital contexts. This finding acquires enhanced theoretical significance when considered alongside methodological innovations employed in this investigation. Traditional survey methodologies lack capacity to detect micro-temporal cognitive dynamics characterizing digital content consumption(Otto et al. 2022). Through computational operationalization of cognitive load via multimodal feature complexity, cross-modal interference patterns, and semantic density variations, previously invisible influence mechanisms emerge. The concentration of predictive power, wherein 42.2% of model importance derives from merely 10.4% of features, suggests mental activation represents not simply another factor but potentially the master mechanism determining whether content translates into behavioral intention. The remarkably weak canonical correlation between audio and visual features (CCA = 0.028) provides crucial corroborating evidence. In contrast to professional content where audiovisual elements achieve careful synchronization, PUGC creators appear to prioritize authenticity over production coherence. This production heterogeneity paradoxically increases cognitive demands, compelling viewers toward active sense-making rather than passive consumption. The implications prove profound, suggesting that amateur aesthetics characteristic of PUGC achieve effectiveness not despite but because of cognitive demands imposed, thereby transforming viewers from spectators into mental participants. 5.2 The Overseas Cultural Experience Paradox and Virtual Travel Substitution Among our findings, the negative association between Overseas Cultural Experience videos and travel intention (β = -4.59, p < 0.001) presents the most theoretically provocative result. This counterintuitive pattern, wherein rich cultural content diminishes rather than stimulates travel desire, demands careful theoretical consideration through multiple explanatory frameworks. The mediated substitution hypothesis advanced by Liu et al. (2023)(Liu et al. 2023) offers one explanatory pathway, proposing that immersive virtual experiences satisfy psychological needs motivating travel, including novelty seeking, cultural learning, and social distinction, without requiring physical displacement. Our data provides supporting evidence through this category's distinctive profile, characterized by highest semantic richness evidenced in CLIP embedding magnitude, lowest Danmaku density (M = 0.73), and moderate cognitive load levels. This combination suggests contemplative consumption patterns wherein viewers achieve cultural gratification through viewing alone. Nevertheless, alternative mechanisms warrant careful consideration. Selection bias potentially explains these patterns if audiences consuming extensive cultural content comprise primarily armchair travelers seeking cultural knowledge without corresponding travel intention. Alternatively, cognitive overload mechanisms, supported by non-linear models revealing threshold effects(Hossain and Yeasin 2017), suggest excessive cultural complexity might overwhelm rather than inspire, particularly when viewers lack appropriate cultural scaffolding for interpretation. Moreover, social proof deficit, evidenced by low Danmaku density, removes collective enthusiasm that typically transforms individual interest into shared aspiration. Most compelling, temporal analysis reveals Overseas Cultural Experience videos demonstrate strongest Granger causality patterns, with visual elements preceding audio responses that subsequently generate textual commentary in 14.3% of analytical windows. This sophisticated orchestration paradoxically may undermine travel intention by providing such comprehensive virtual experiences that physical travel appears redundant. 5.3 Non-linear Dynamics and Optimal Cognitive Load Configurations The stark performance disparity between linear models (R² = 0.033) and non-linear approaches (R² = 0.378, optimized) provides empirical validation for threshold-based influence mechanisms. SHAP analysis reveals specific non-linearities wherein cognitive load demonstrates inverted-U relationships with travel intention, achieving peak influence at moderate complexity levels (standardized value ≈ 0.4) before subsequent decline. While this pattern aligns with flow theory conceptualized by Csikszentmihalyi (1990)(Mihaly Csikszentmihalyi 1990), manifestation differs substantially between tourism contexts and traditional learning environments. Category-specific optimal zones discovered through systematic analysis provide actionable insights for content optimization. Destination Attraction content maximizes influence at lower cognitive load levels (0.2–0.3 standardized units) combined with elevated social density (1.50 comments per window), suggesting viewers seek social validation for aesthetic experiences. Conversely, Tourism Narrative Communication achieves optimal performance at higher cognitive load levels (0.5–0.6 standardized units), indicating audiences expect and reward storytelling complexity. Meanwhile, Overseas Cultural Experience demonstrates the flattest response curve, suggesting cognitive load may prove less determinative than cultural authenticity markers for this content category. These patterns necessitate fundamental reconceptualization of digital tourism influence. Rather than maximizing production quality or minimizing cognitive effort, effective PUGC operates within category-specific cognitive configurations where challenge activates without overwhelming processing capacity. 5.4 Methodological Innovation and Theoretical Advancement Integration Our computational approach transcends mere efficiency gains in measuring established phenomena, instead revealing influence mechanisms invisible to traditional methodologies. Extraction of 102 objective features capturing millisecond-level multimodal dynamics enabled discovery of three critical phenomena previously unobservable. First, attention cascades demonstrate predictable operation across modalities, wherein visual changes trigger audio responses that subsequently generate Danmaku commentary. This temporal choreography, detected through Granger causality analysis, suggests successful creators intuitively orchestrate influence sequences that traditional cross-sectional analysis(Spector 2019) cannot detect. Second, cognitive interference patterns between modalities reveal conditions under which information streams compete versus complement. The negative temporal consistency coefficient (β = -5.20) indicates monotony diminishes engagement, whereas controlled variation maintains cognitive activation, a finding achievable only through computational temporal analysis. Third, collective sense-making dynamics emerge through Danmaku density bursts at specific narrative moments, revealing transformation points where individual viewing becomes social experience. These crystallization points, characterized by concentrated travel intention expressions, provide blueprints for engineering viral influence moments. 5.5 Tourism Marketing Transformation from Broadcasting to Cognitive Choreography Practical implications emerging from our findings necessitate fundamental reconsideration of tourism marketing strategies. Traditional paradigms emphasizing inspiration through beauty, information through features, and persuasion through benefits assume passive viewers processing additive information. Conversely, our evidence indicates active viewers navigate complex information ecosystems wherein cognitive engagement determines influence outcomes. Destination marketing organizations must therefore shift emphasis from production quality toward cognitive design. Rather than maximizing visual appeal, content should optimize cognitive load within category-specific configurations(Bai et al. 2025). Destination attraction content should maintain visual stability while maximizing social interaction opportunities(Malafouris et al. 2021). Tourism narratives require strategic multimodal variation sustaining engagement across extended viewing periods(Mattei 2023). Cultural content must carefully balance immersion with aspiration, avoiding substitution effects wherein virtual satisfaction replaces travel motivation(Hu et al. 2023). Content creators receive validation for PUGC approaches while gaining optimization guidance. Weak audiovisual correlation (CCA = 0.028) suggests production perfection may prove counterproductive, as authentic heterogeneity engages more effectively than polished coherence. Creators should therefore focus on cognitive choreography through strategic information revelation, controlled complexity escalation, and social interaction facilitation(Sigala 2016). Platform designers encounter algorithmic innovation opportunities suggested by our results. Current recommendation systems optimizing view duration or engagement metrics may overlook cognitive activation patterns driving behavioral intention(Wen et al. 2019). Platforms could develop creator tools facilitating attention cascade orchestration, optimal Danmaku density maintenance, and cognitive load monitoring capabilities(He and Tang 2017; Yang 2020b; Peng et al. 2025). 6. Limitations and Future Research Directions While this investigation advances understanding of digital tourism influence mechanisms, several methodological and theoretical boundaries illuminate both constraints and opportunities for future inquiry. Our reliance on comment-based intention measures captures expressed enthusiasm among 8.7% of viewers rather than actual booking behaviors, representing the most fundamental limitation. Paradoxically, this constraint strengthens theoretical contributions, as cognitive mechanisms dominating even expressed intention likely play stronger roles in actual behavior requiring greater commitment. Future research should establish conversion pathways through platform-DMO partnerships enabling behavioral validation. Similarly, single-platform data from Bilibili constrains cross-cultural generalizability while providing exceptional ecosystem depth. Bilibili's distinctive Danmaku feature, absent from Western platforms, enabled unique insights into synchronized social viewing. Comparative analysis across YouTube, TikTok, and Instagram Reels would distinguish universal cognitive mechanisms from platform-specific affordances, particularly revealing whether platforms lacking Danmaku exhibit stronger audiovisual effects due to reduced attentional competition. The substantial overfitting observed in certain models, with train-test gaps reaching 0.430, raises critical questions about cognitive feature stability. Rather than purely methodological weakness, this instability may carry theoretical significance, suggesting cognitive influence mechanisms are inherently context-dependent and require adaptive rather than universal modeling approaches. Whether overfitting reflects genuine cognitive complexity or measurement limitations remains an empirical question demanding resolution. These limitations nonetheless establish several theoretical frontiers warranting systematic investigation. The cognitive efficiency paradox, wherein five cognitive features outperform 22 audiovisual features, suggests information compression mechanisms extract essential decision inputs through either evolutionary spatial adaptations or learned digital heuristics. Equally intriguing, the authenticity-chaos hypothesis emerges from remarkably weak audiovisual correlation (CCA = 0.028), contradicting media richness theory while suggesting production incoherence paradoxically signals authenticity. Experimental manipulation of production coherence could establish causal relationships and boundary conditions for this counterintuitive phenomenon. Most provocatively, the substitution threshold model implied by Overseas Cultural Experience findings indicates certain content satisfies rather than stimulates travel desire. As virtual reality tourism emerges, identifying precise thresholds between inspiration and substitution becomes critical for industry sustainability. Research mapping psychological needs satisfied through virtual versus physical travel could inform strategic responses to technological disruption. Meanwhile, the collective intelligence mechanism manifested through Danmaku raises fundamental questions about social information processing. Eye-tracking studies could determine whether viewers process synchronized commentary as informational input or atmospheric enhancement, distinguishing wisdom of crowds from conformity pressure. Our findings arrive at a critical juncture as tourism marketing confronts AI-generated content, virtual reality experiences, and intensifying attention economy dynamics. Three implications emerge with particular force. First, cognitive engagement may constitute the emerging scarce resource in tourism marketing, as infinite AI-generated content confronts finite human processing capacity. Second, platform mediation will intensify given evidence that affordances fundamentally alter influence mechanisms, necessitating platform-specific rather than universal strategies. Third, boundaries between virtual and physical tourism will progressively blur, with our Overseas Cultural Experience findings previewing direct competition between mediated and embodied experiences. Tourism's value proposition must therefore evolve beyond visual consumption toward experiences that resist virtualization. These considerations suggest tourism marketing stands at an inflection point where traditional paradigms yield to cognitive-interactive frameworks. Success will increasingly depend on understanding and orchestrating mental engagement rather than maximizing sensory stimulation, a transformation our findings both document and accelerate. 7. Conclusion This investigation illuminates a fundamental reorientation in digital tourism influence, wherein cognitive activation supersedes sensory persuasion as the primary mechanism driving travel intention formation. Computational analysis of 2,650 PUGC videos from Bilibili reveals that cognitive engagement accounts for substantially greater variance in behavioral intention than audiovisual attributes, challenging orthodox assumptions underlying tourism marketing practice. Three theoretical contributions emerge from this analysis. First, the identification of non-linear, threshold-based influence patterns explains the persistent failure of linear frameworks in digital contexts, demonstrating that cognitive features generate predictive power through discontinuous rather than additive relationships. Second, Danmaku's transformation of solitary viewing into collective sense-making constitutes a novel social influence pathway, wherein synchronized commentary creates cognitive scaffolding absent from traditional media consumption. Third, the counterintuitive finding that culturally rich content diminishes rather than enhances travel intention illuminates an emerging substitution dynamic, presaging fundamental challenges as virtual experiences increasingly approximate physical travel's psychological rewards. Methodologically, this study validates computational approaches for detecting influence mechanisms beyond the reach of conventional methods. Extraction and analysis of 102 objective features capturing millisecond-level multimodal dynamics revealed attention cascades, cognitive interference patterns, and collective crystallization moments invisible to survey-based inquiry. Such granular temporal resolution proves essential for understanding how influence unfolds through orchestrated sequences rather than static attributes. These findings acquire heightened significance as tourism marketing navigates technological disruption through AI-generated content and immersive virtual experiences. The primacy of cognitive engagement over production quality suggests competitive advantage will increasingly derive from mental activation rather than sensory sophistication. Content creators who master cognitive choreography—the strategic orchestration of complexity, variation, and social interaction—will shape travel decisions more effectively than those pursuing traditional aesthetic excellence. Thus, the pathway from watching to wishing, this investigation reveals, traverses not the eye but the mind, positioning cognitive design as tourism marketing's emerging imperative. Declarations Author Contribution F.Y. designed the study, developed the methodology, implemented software, conducted the investigation, and prepared the original manuscript draft.M.Y. performed data curation, validation, and contributed to reviewing and editing the manuscript.S.S. provided supervision and contributed to the critical review of the manuscript.X.W. provided supervision and project oversight and contributed to manuscript revision.All authors reviewed and approved the final manuscript. References Alaei AR, Becken S, Stantic B (2019) Sentiment Analysis in Tourism: Capitalizing on Big Data. Journal of Travel Research 58:175–191. https://doi.org/10.1177/0047287517747753 Anh HH (2024) Harmonizing Engagement: The Impact of Music on Consumer Interaction in Social Media Marketing. IJEBMR 08:182–202. https://doi.org/10.51505/IJEBMR.2024.81013 Asyraff MA, Hanafiah M, Aminuddin M, Mahdzar N (2023) Adoption of the stimulus-organism-response (S-O-R) model in hospitality and tourism research: systematic literature review and future research directions adoption of the stimulus-organism-response (S-O-R) model in hospitality and tourism research: systematic literature review and future research directions. Asia-pac J Innov Hosp Tour (APJIHT) 12:2023 B. Joseph Pine II, Gilmore James H. (1998) Welcome to the Experience Economy. Harvard Business Review Bai S, Li Z, He H, Fan W (2025) Engaging tourists from city aesthetics: evidence from multi-modal analysis of computer vision and text mining. APJML. https://doi.org/10.1108/APJML-02-2025-0287 Bailenson J, Guadagno R, Aharoni E, et al (2004) Comparing behavioral and self-report measures of embodied agents’ social presence in immersive virtual environments Baltrusaitis T, Ahuja C, Morency L-P (2019) Multimodal Machine Learning: A Survey and Taxonomy. IEEE Trans Pattern Anal Mach Intell 41:423–443. https://doi.org/10.1109/TPAMI.2018.2798607 Bayerl PS, Paul KI (2011) What Determines Inter-Coder Agreement in Manual Annotations? A Meta-Analytic Investigation. Computational Linguistics 37:699–725. https://doi.org/10.1162/COLI_a_00074 (Brian) Lin M-T, Zhu D, Liu C, Kim PB (2022) A systematic review of empirical studies of pro-environmental behavior in hospitality and tourism contexts. IJCHM 34:3982–4006. https://doi.org/10.1108/IJCHM-12-2021-1478 Brunken R, Plass JL, Leutner D (2003) Direct Measurement of Cognitive Load in Multimedia Learning. Educational Psychologist 38:53–61. https://doi.org/10.1207/S15326985EP3801_7 Buckingham D (2009) A Commonplace Art? Understanding Amateur Media Production. Video Cultures 23–50. https://doi.org/10.1057/9780230244696_2 Cauchard JR, Gover W, Chen WL, et al (2024) Orality, multimodality and creativity in digital writing: Chinese users’ experiences and practices with bullet comments on Bilibili . Social Semiotics 34:368–394. https://doi.org/10.1080/10350330.2022.2120387 Caulley DN (1994) Review & Booknote: The Unobtrusive Researcher: A Guide to Methods. Media Information Australia 73:118–118. https://doi.org/10.1177/1329878X9407300133 Chemin M, Silva CP, Vikou SVP (2025) User-generated content (UGC) in tourist attractions and destinations: systematic literature review and perspectives for management Chen M (1986) The Impact of Authenticity and Credibility Factors on Consumer Behavior in the Sustainable Fashion Industry on Weibo–Taking Micro-Influencers and Mega-Influencers as Examples. Tourism Management 91:322–328. https://doi.org/10.54254/2754-1169/91/20241077 Chen Y, Gao Q, Rau P-LP (2015) Understanding Gratifications of Watching Danmaku Videos – Videos with Overlaid Comments. In: Rau PLP (ed) Lecture Notes in Computer Science. Springer International Publishing, Cham, pp 153–163 Chen Z (2025) Theoretical development of the tourist experience: a future perspective. Tourism Recreation Research 50:199–213. https://doi.org/10.1080/02508281.2023.2255939 Cialdini R (1993) Influence: Science and Practice Diehl M, Wahl H-W, Freund A (2017) Ecological Validity as a Key Feature of External Validity in Research on Human Development. Research in Human Development 14:177–181. https://doi.org/10.1080/15427609.2017.1340053 Dimitriou CK, AbouElgheit E (2019) Understanding generation Z’s travel social decision-making. Tour hosp manag 25:311–334. https://doi.org/10.20867/thm.25.2.4 Ehlers J (2020) Exploring the effect of transient cognitive load on bodily arousal and secondary task performance. In: Proceedings of Mensch und Computer 2020. ACM, New York, NY, USA, pp 7–10 Fan M, Ostic D, Han K, Shar S (2023) The Impact of Danmaku Information Quality on Consumers’ Impulsive Consumption Behavior. Proceedings 2023:10474. https://doi.org/10.5465/AMPROC.2023.10474abstract Fang J, Chen L, Wen C, Prybutok VR (2018) Co-viewing Experience in Video Websites: The Effect of Social Presence on E-Loyalty. International Journal of Electronic Commerce 22:446–476. https://doi.org/10.1080/10864415.2018.1462929 Fazli-Salehi R, Jahangard M, Torres IM, et al (2022) Social media reviewing channels: the role of channel interactivity and vloggers’ self-disclosure in consumers’ parasocial interaction. JCM 39:242–253. https://doi.org/10.1108/JCM-06-2020-3866 Filieri R, Christodoulides G, Nicolau JL (2025) Emerging Sources, Formats, Channels, Devices, and Audiences in Modern eWOM Communications. Psychology and Marketing 42:2430–2443. https://doi.org/10.1002/mar.22239 Fotis J, Buhalis D, Rossides N (2012) Social Media Use and Impact during the Holiday Travel Planning Process. Information and Communication Technologies in Tourism 2012 13–24. https://doi.org/10.1007/978-3-7091-1142-0_2 Gong W, Li X (2017) Engaging fans on microblog: the synthetic influence of parasocial interaction and source characteristics on celebrity endorsement Gretzel U (2018) Influencer Marketing in Travel and Tourism Grimmer J, Stewart BM (2013) Text as Data: The Promise and Pitfalls of Automatic Content Analysis Methods for Political Texts. Polit anal 21:267–297. https://doi.org/10.1093/pan/mps028 Grootendorst M (2020) BERTopic: Neural topic modeling with a class-based TF-IDF procedure Haji FA, Rojas D, Childs R, et al (2015) Measuring cognitive load: performance, mental effort and simulation task complexity. Medical Education 49:815–827. https://doi.org/10.1111/medu.12773 Hatfield E, Cacioppo JT, Rapson RL (1993) Emotional Contagion. Curr Dir Psychol Sci 2:96–100. https://doi.org/10.1111/1467-8721.ep10770953 He K, Zhang X, Ren S, Sun J (2016) Deep Residual Learning for Image Recognition He Y, Tang TY (2017) Recommending highlights in Anime movies: Mining the real-time user comments “DanMaKu.” In: 2017 Intelligent Systems Conference (IntelliSys). pp 319–322 Horton D, Richard Wohl R (1956) Mass Communication and Para-Social Interaction. Psychiatry 19:215–229. https://doi.org/10.1080/00332747.1956.11023049 Hossain G, Yeasin M (2017) Analysis of Cognitive Dissonance and Overload through Ability-Demand Gap Models. IEEE Trans Cogn Dev Syst 9:170–182. https://doi.org/10.1109/TAMD.2015.2450681 Hu Y, Phawitpiriyakliti C, Terason S (2023) Tourist Satisfaction in Virtual Reality Immersive Experiences: Implications for the Tourism Industry. The Journal of Pacific Institute of Management Science (Humanities and Social Science) 9:115–128 Huang Z (2024) Study on the Optimization of Bilibili’s Business Model —— Based on the Business Model Canvas. HBEM 36:501–508. https://doi.org/10.54097/z6hzst93 Jun SH, Vogt C (2013) TRAVEL INFORMATION PROCESSING APPLYING A DUAL-PROCESS MODEL. Annals of Tourism Research 40:191–212. https://doi.org/10.1016/j.annals.2012.09.001 KAMEI K, TOYOTA A, KUSHIDA J (2012) Upsurge of Viewer’s Emotion by Video Sharing Using Pseudo Synchronization. J SOFT 24:944–953. https://doi.org/10.3156/jsoft.24.944 Kan Y (2023) Research on the Integration of Multimodal Large Language Models (MLLM) and Augmented Reality (AR) for Smart Navigation with Real-Time Cross-Language Interaction and Cognitive Load Balancing Strategies. Journal of Business Research 5:2540752. https://doi.org/10.1142/S0129156425407521 Kim AJ, Ko E (2012) Do social media marketing activities enhance customer equity? An empirical study of luxury fashion brand. Journal of Business Research 65:1480–1486. https://doi.org/10.1016/j.jbusres.2011.10.014 Larson M, Melenhorst M, Menéndez M, Xu P (2014) Using crowdsourcing to capture complexity in human interpretations of multimedia content. In: Ionescu B, Benois-Pineau J, Piatrik T, Quénot G (eds) Fusion in Computer Vision: Understanding Complex Visual Content. Springer International Publishing, Cham, pp 229–269 Leng J, Zhu J, Wang X, Gu X (2016) Identifying the Potential of Danmaku Video from Eye Gaze Data. In: 2016 IEEE 16th International Conference on Advanced Learning Technologies (ICALT). pp 288–292 Leung D, Law R, van Hoof H, Buhalis D (2013) Social Media in Tourism and Hospitality: A Literature Review. Journal of Travel & Tourism Marketing 30:3–22. https://doi.org/10.1080/10548408.2013.750919 Li F, Wang W, Lai W, et al (2025) Unveiling the Multidimensional Nature of the Intention–Behavior Gap. European Journal of Health Psychology 32:34–50. https://doi.org/10.1027/2512-8442/a000162 Li H, Tu X (2024a) Who generates your video ads? The matching effect of short-form video sources and destination types on visit intention. APJML 36:660–677. https://doi.org/10.1108/APJML-04-2023-0300 Li H, Tu X (2024b) Who generates your video ads? The matching effect of short-form video sources and destination types on visit intention. APJML 36:660–677. https://doi.org/10.1108/apjml-04-2023-0300 Li S, Scott N, Walters G (2015) Current and potential methods for measuring emotion in tourism experiences: a review. Current Issues in Tourism 18:805–827. https://doi.org/10.1080/13683500.2014.975679 Liu C (2024) A Study on the Correlation between Danmaku Videos and Audiences’ Repetitive Viewing Behaviours - A Case Study of Bilibili Danmaku Video Network Videos. HC 1:. https://doi.org/10.61173/knqn8g59 Liu Q, Xu L, Feng W, et al (2023) Is tourism live streaming a double-edged sword? The paradoxical impact of online flow experience on travel intentions. Journal of Travel & Tourism Marketing 40:744–763. https://doi.org/10.1080/10548408.2023.2293016 Liu W, Wang Z, Jian L, Sun Z (2024) How broadcasters’ characteristics affect viewers’ loyalty: the role of parasocial relationships Luo J (1986) The Influence of the Credibility of Brand Content on Social Media Platforms on Consumers’ Purchasing Decisions and Its Communication Mechanism. Tourism Management 1:. https://doi.org/10.61173/9g99dd30 Luo J, Chen M, Chen Y, et al (2021) Understanding and managing the threat of common method bias: Detection, prevention and control. Tourism Management 86:104330. https://doi.org/10.1016/j.tourman.2021.104330 Lyu B (2021) How is the Purchase Intention of Consumers Affected in the Environment of E-commerce Live Streaming? Malafouris L, Gosden C, Masson M, et al (2021) Building Social Media Engagement on Instagram by Using Visual Aesthetics and Message Orientation Strategy: A Content Analysis on Instagram Content of Indonesia Tourism Destinations. JICP 4:129–138. https://doi.org/10.32535/jicp.v4i3.1304 Mattei E (2023) Multimodal corpus analysis of digital tourism narratives: A data-driven approach based on Systemic Functional Linguistics and social semiotics Mehra P (2023) Unexpected surprise: Emotion analysis and aspect based sentiment analysis (ABSA) of user generated comments to study behavioral intentions of tourists. Tourism Management Perspectives 45:101063. https://doi.org/10.1016/j.tmp.2022.101063 Mehrabian A, Russell JA (1980) An Approach to Environmental Psychology. MIT Press, Cambridge, MA, USA Mihaly Csikszentmihalyi (1990) Flow: The Psychology of Optimal Experience Naab TK, Karnowski V, Schlütz D, et al (2019) Reporting Mobile Social Media Use: How Survey and Experience Sampling Measures Differ. Adaptive Behavior 13:126–147. https://doi.org/10.1080/19312458.2018.1555799 Nguyen PMB, Pham LX, Tran DK, Truong GNT (2024) A systematic literature review on travel planning through user-generated video. Journal of Vacation Marketing 30:553–581. https://doi.org/10.1177/13567667231152935 Otto LP, Thomas F, Glogger I, De Vreese CH (2022) Linking Media Content and Survey Data in a Dynamic and Digital Media Environment – Mobile Longitudinal Linkage Analysis. Digital Journalism 10:200–215. https://doi.org/10.1080/21670811.2021.1890169 Pan X (2023) Motivations for game stream spectatorship: A content analysis of Danmaku on Bilibili. Global Media and China 8:190–212. https://doi.org/10.1177/20594364231179750 Pascual-Fraile M del P, Villacé-Molinero T, Talón-Ballestero P, Chaperon S (2025) Post-pandemic collaborative destination marketing: Effectiveness and impact on different generational audiences. Front Hum Neurosci 31:665–682. https://doi.org/10.1177/13567667231224091 Peinado O, Shim M (2024) The intersection of “real” and “reel”: An investigation of K-pop idol dual self-presentation, paid advertisements, and fan engagement. Computers in Human Behavior 161:108414. https://doi.org/10.1016/j.chb.2024.108414 Peinado O, Shim M, Alam A, et al (2020) Video Big Data Analytics in the Cloud: A Reference Architecture, Survey, Opportunities, and Open Research Issues. IEEE Access 8:152377–152422. https://doi.org/10.1109/ACCESS.2020.3017135 Peng G, Wang X, Li J, Wu J (2025) Is it beneficial for consumers to ask more questions in Danmaku? The inverted U-shaped effect of information-seeking Danmaku density on e-commerce livestream sales. Electron Commer Res. https://doi.org/10.1007/s10660-025-09959-1 Pertiwi E, Sanusi AP (2023) Storytelling in the Digital Age: Examining the Role and Effectiveness in Communication Strategies of Social Media Content Creators. Palakka Media Islam Commun 4:25–34. https://doi.org/10.30863/palakka.v4i1.5082 Pham MT (2007) Emotion and Rationality: A Critical Review and Interpretation of Empirical Evidence. Review of General Psychology 11:155–178. https://doi.org/10.1037/1089-2680.11.2.155 Pink A, Newton PM (2020) Decorative animations impair recall and are a source of extraneous cognitive load. Advances in Physiology Education 44:376–382. https://doi.org/10.1152/advan.00102.2019 Podsakoff PM, Organ DW (1986) Self-Reports in Organizational Research: Problems and Prospects. Journal of Management 12:531–544. https://doi.org/10.1177/014920638601200408 Polat E, Çelik F, Ibrahim B, Köseoglu MA (2023) Unpacking the power of user-generated videos in hospitality and tourism: a systematic literature review and future direction. Journal of Travel & Tourism Marketing 40:894–914. https://doi.org/10.1080/10548408.2023.2296655 Pop R-A, Săplăcan Z, Dabija D-C, Alt M-A (2022) The impact of social media influencers on travel decisions: the role of trust in consumer decision journey. Taylor & Francis Radford A, Kim JW, Hallacy C, et al (2003) Learning Transferable Visual Models From Natural Language Supervision Rahwan I, Cebrian M, Obradovich N, et al (2024) Danmaku consistency reduces consumer purchases during live streaming: A dual-process model. Psychology and Marketing 41:2591–2607. https://doi.org/10.1002/mar.22074 Ross P, Paas F, Tuovinen JE, et al (2011) Is there an expertise of production? The case of new media producers. New Media & Society 13:912–928. https://doi.org/10.1177/1461444810385393 Ross P, Paas F, Tuovinen JE, et al (2003) Cognitive Load Measurement as a Means to Advance Cognitive Load Theory. Educational Psychologist 38:63–71. https://doi.org/10.1207/S15326985EP3801_8 Samson A, Voyer BG (2012) Two minds, three ways: dual system and dual process models in consumer psychology. AMS Rev 2:48–71. https://doi.org/10.1007/s13162-012-0030-9 Schwenzow J, Hartmann J, Schikowsky A, Heitmann M (2021) Understanding videos at scale: How to extract insights for business research. Journal of Business Research 123:367–379. https://doi.org/10.1016/j.jbusres.2020.09.059 Shah RR, Mahata D, Choudhary V, Bajpai R (2018) Multimodal Semantics and Affective Computing from Multimedia Content. Advances in Multimedia and Interactive Technologies 359–382. https://doi.org/10.4018/978-1-5225-5246-8.ch014 Sigala M (2016) Social Media and the Co-creation of Tourism Experiences. The Handbook of Managing and Marketing Tourism Experiences 85–111. https://doi.org/10.1108/978-1-78635-290-320161033 Silaban PH, Chen W-K, Silaban BE, et al (2023) Demystifying Tourists’ Intention to Visit Destination on Travel Vlogs: Findings from PLS-SEM and fsQCA. Emerg Sci J 7:867–889. https://doi.org/10.28991/ESJ-2023-07-03-015 Skulmowski A, Rey GD (2015) Measuring Cognitive Load in Embodied Learning Settings. Front Psychol 8:815–827. https://doi.org/10.3389/fpsyg.2017.01191 Spector PE (2019) Do Not Cross Me: Optimizing the Use of Cross-Sectional Designs. J Bus Psychol 34:125–137. https://doi.org/10.1007/s10869-018-09613-8 Stone AA, Shiffman S (2002) Capturing momentary, self-report data: A proposal for reporting guidelines. ann behav med 24:236–243. https://doi.org/10.1207/S15324796ABM2403_09 SU Y, LIU M, ZHAO N, et al (2021) Identifying psychological indexes based on social media data: A machine learning method. Adv Psychol Sci 29:571–585. https://doi.org/10.3724/SP.J.1042.2021.00571 Sweller J (1988) Cognitive Load During Problem Solving: Effects on Learning. Cognitive Science 12:257–285. https://doi.org/10.1207/s15516709cog1202_4 Tan-intaraarj P (2024) Exploring influence attempts, wishful identification, parasocial relationships, and behavioral loyalty among Thai game live-streamers and their viewers. Asian Journal of Communication 34:178–194. https://doi.org/10.1080/01292986.2024.2315580 Thein T, Westbrook RF, Harris JA (2022) How the associative strengths of stimuli combine in compound: Summation and overshadowing. Journal of Experimental Psychology: Animal Behavior Processes 34:155–166. https://doi.org/10.1037/0097-7403.34.1.155 Turner J, O’Leary M (2012) Targets’ practices: how people allocate their attention among multiple streams of incoming information Urry J (1990) The Tourist Gaze: Leisure and Travel in Contemporary Societies, 第 1st 版. SAGE Publications Ltd, London Vaswani A, Shazeer N, Parmar N, et al (2023) Attention Is All You Need Ward L, Glancy M, Bowman S, Armstrong M (2020) The impact of new forms of media on production tools and practices Wen H, Yang L, Estrin D (2019) Leveraging post-click feedback for content recommendations. In: Proceedings of the 13th ACM Conference on Recommender Systems. Association for Computing Machinery, New York, NY, USA, pp 278–286 Wu G, Pei X, Wang D, et al (2024) I Bond, I Engage, I Visit: Investigating the Effects of Vloggers Tourist Engagement and Its Outcome on Tourist Attitudes. Journal of Travel Research. https://doi.org/10.1177/00472875241276546 Wu Q, Sang Y, Huang Y (2019) Danmaku: A New Paradigm of Social Interaction via Online Videos. Trans Soc Comput 2:1–24. https://doi.org/10.1145/3329485 Wu S, Rizoiu M-A, Xie L (2018) Beyond Views: Measuring and Predicting Engagement in Online Videos. ICWSM 12:. https://doi.org/10.1609/icwsm.v12i1.15031 Wyer RS, Kardes FR (2020) A Multistage, Multiprocess Analysis of Consumer Judgment: A Selective Review and Conceptual Framework. J Consum Psychol 30:339–364. https://doi.org/10.1002/jcpy.1158 Xu P, Zhu X, Clifton DA (2023) Multimodal Learning With Transformers: A Survey. IEEE Trans Pattern Anal Mach Intell 45:12113–12132. https://doi.org/10.1109/TPAMI.2023.3275156 Yang J, Zeng Y, Liu X, Li Z (2022) Nudging interactive cocreation behaviors in live-streaming travel commerce: The visualization of real-time danmaku. Journal of Hospitality and Tourism Management 52:184–197. https://doi.org/10.1016/j.jhtm.2022.06.015 Yang Y (2020a) The danmaku interface on Bilibili and the recontextualised translation practice: a semiotic technology perspective. Social Semiotics 30:254–273. https://doi.org/10.1080/10350330.2019.1630962 Yang Y (2020b) The danmaku interface on Bilibili and the recontextualised translation practice: a semiotic technology perspective. Taylor & Francis Yu Y, Huang S, Liu Y, Tan Y (2025) Emotions in Online Content Diffusion. Information Systems Research. https://doi.org/10.1287/isre.2022.0611 Yuan H, Lu K, Ausaf A, Zhu M (2024) Constant or inconstant? The time-varying effect of danmaku on user engagement in online video platforms. INTR 35:771–797. https://doi.org/10.1108/INTR-06-2023-0479 Yüksel A (2017) A critique of “Response Bias” in the tourism, travel and hospitality research. Tourism Management 59:376–384. https://doi.org/10.1016/j.tourman.2016.08.003 Zhou J, Zhou J, Ding Y, Wang H (2018) The Magic of Danmaku: A Social Interaction Perspective of Gift Sending on Live Streaming Platforms. SSRN Journal. http://dx.doi.org/10.2139/ssrn.3289119 Zhu S, Zhu X, Yao Y, Cheong CM (2025) Profiling the differences in strategy use in online multimodal reading: Associations with self-efficacy and reading task performance. Studies in Educational Evaluation 87:101507. https://doi.org/10.1016/j.stueduc.2025.101507 Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 06 Jan, 2026 Reviews received at journal 29 Dec, 2025 Reviews received at journal 16 Dec, 2025 Reviewers agreed at journal 08 Dec, 2025 Reviewers agreed at journal 02 Dec, 2025 Reviewers invited by journal 01 Dec, 2025 Editor assigned by journal 24 Nov, 2025 Submission checks completed at journal 19 Nov, 2025 First submitted to journal 18 Nov, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8141450","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":553934638,"identity":"0abfea1e-6df0-4008-aa0c-f23397f0ac79","order_by":0,"name":"Feng Ye","email":"","orcid":"","institution":"Communication University of Zhejiang","correspondingAuthor":false,"prefix":"","firstName":"Feng","middleName":"","lastName":"Ye","suffix":""},{"id":553934639,"identity":"b0aca889-8f95-4446-9d8e-97eb429141a6","order_by":1,"name":"Min Yin","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAu0lEQVRIiWNgGAWjYJCCAwwVDMwghgQJWs6QqoWBsQ1CE6eFf9oZw8OF8+rYDQ4wH7zNw2CXR1CLxO0cg8Mzt7ExGxxgS7bmYUguJmwNSAvvNh6gFh4zaR6GA4kNhHTIg7XMkQBq4f9GnBYDsJYGA5AtbMRpMbydVnCY51gCs+RhNmPLOQbJhLXI3U7e/Jmnpi6Z73jzwxtvKuwIa2Fg4DAAkcmQyDQgrB4I2B+ASDui1I6CUTAKRsHIBABEtzb4vLE2kgAAAABJRU5ErkJggg==","orcid":"","institution":"Zhejiang University","correspondingAuthor":true,"prefix":"","firstName":"Min","middleName":"","lastName":"Yin","suffix":""},{"id":553934640,"identity":"f4a352bc-4f4f-4098-b998-511bbbae2bfe","order_by":2,"name":"Shouqian Sun","email":"","orcid":"","institution":"Zhejiang University","correspondingAuthor":false,"prefix":"","firstName":"Shouqian","middleName":"","lastName":"Sun","suffix":""},{"id":553934641,"identity":"967a5b31-1460-4d05-a9ed-58126500c03b","order_by":3,"name":"Xuanzheng Wang","email":"","orcid":"","institution":"Central Academy of Fine Arts","correspondingAuthor":false,"prefix":"","firstName":"Xuanzheng","middleName":"","lastName":"Wang","suffix":""}],"badges":[],"createdAt":"2025-11-18 06:08:14","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8141450/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8141450/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":97666501,"identity":"ca502e02-13fa-4788-972a-5bfe8f9fcd37","added_by":"auto","created_at":"2025-12-08 09:21:22","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":2110792,"visible":true,"origin":"","legend":"","description":"","filename":"FromWatchingtoWishingannoymous.docx","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/aae0bd13226e86973fbe1d45.docx"},{"id":97667195,"identity":"0dd3ae0c-8c74-47b0-a133-b771c7786a45","added_by":"auto","created_at":"2025-12-08 09:23:00","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":6183,"visible":true,"origin":"","legend":"","description":"","filename":"2d242c96728f444e9e03fed4b338f07b.json","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/3b029673968879c5c987299f.json"},{"id":97666459,"identity":"ee19a963-5b2c-4c94-8e05-29a130f0e91b","added_by":"auto","created_at":"2025-12-08 09:21:15","extension":"xml","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":188563,"visible":true,"origin":"","legend":"","description":"","filename":"2d242c96728f444e9e03fed4b338f07b1enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/080624e24b82434b32afe8ef.xml"},{"id":97666316,"identity":"79a1443d-ad37-4e61-8fb8-aa90ed46d7b9","added_by":"auto","created_at":"2025-12-08 09:21:00","extension":"png","order_by":27,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":48776,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/5209a5c293f20a3025c9114e.png"},{"id":97666201,"identity":"78ca7c4c-383e-4028-a387-7c8ac20d40e7","added_by":"auto","created_at":"2025-12-08 09:20:37","extension":"png","order_by":28,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":58871,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage10.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/691855af22b1c403681bd48d.png"},{"id":97666481,"identity":"e12b97b2-fa8b-4107-98ef-639e45cf6344","added_by":"auto","created_at":"2025-12-08 09:21:19","extension":"png","order_by":29,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":59494,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage11.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/2c413bcc212ad70521c0d851.png"},{"id":97666818,"identity":"0c390254-05c9-4332-8c46-2483a4dff7a5","added_by":"auto","created_at":"2025-12-08 09:22:11","extension":"png","order_by":30,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":31759,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage12.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/9b3151fcbf4c489406ba2b35.png"},{"id":97666241,"identity":"63a70439-2d49-4d90-9e11-7306a49d5b7a","added_by":"auto","created_at":"2025-12-08 09:20:42","extension":"png","order_by":31,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":48776,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/b3fe6fd4711ba35dff7dc775.png"},{"id":97666882,"identity":"081453b3-e25b-48d6-97ca-930e765b942e","added_by":"auto","created_at":"2025-12-08 09:22:18","extension":"png","order_by":32,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":44302,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage14.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/df889ac70bd283bd3b040f2a.png"},{"id":97417103,"identity":"7d148072-42ea-48ad-bdeb-5bfe596059d4","added_by":"auto","created_at":"2025-12-04 07:21:47","extension":"png","order_by":33,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":18322,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage15.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/42232ab66d6607aaa6b9c0f2.png"},{"id":97417113,"identity":"e3995e4e-d8ac-4808-959e-ab8b7c2e9a14","added_by":"auto","created_at":"2025-12-04 07:21:48","extension":"png","order_by":34,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":20811,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage16.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/ef2a0f8c6e6bc9c2723fc8fe.png"},{"id":97417107,"identity":"51239f0a-b64e-434a-8db6-7c9c25a1a266","added_by":"auto","created_at":"2025-12-04 07:21:48","extension":"png","order_by":35,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":9614,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage17.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/8daa4519c932d8d97e9d3fb5.png"},{"id":97417105,"identity":"5685dcf8-adf4-4fdc-bf21-447b648da9eb","added_by":"auto","created_at":"2025-12-04 07:21:47","extension":"png","order_by":36,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":33432,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage18.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/19d6af66ca0f7dca3739e1f5.png"},{"id":97417106,"identity":"6198794c-f848-49ea-ab9d-06356799f371","added_by":"auto","created_at":"2025-12-04 07:21:48","extension":"png","order_by":37,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":22122,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage19.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/e91549a9f76928ff30a9bfd9.png"},{"id":97666440,"identity":"9a8dbecd-4d6f-4b36-bf2a-79eeff550aec","added_by":"auto","created_at":"2025-12-08 09:21:14","extension":"png","order_by":38,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":44302,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage14.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/e09d73d856c0f5e634db5b2b.png"},{"id":97666526,"identity":"09425756-6b47-438f-aae3-992d16520b70","added_by":"auto","created_at":"2025-12-08 09:21:28","extension":"png","order_by":39,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":43135,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage20.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/b879510255c8bc996956b773.png"},{"id":97666571,"identity":"443ba5b7-4d33-4f2c-bf96-8bdebc75df6a","added_by":"auto","created_at":"2025-12-08 09:21:36","extension":"png","order_by":40,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":37520,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage21.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/da7ddaee9efcac232922f5e9.png"},{"id":97417122,"identity":"2a9088d1-9d15-4889-96c1-184537c2603d","added_by":"auto","created_at":"2025-12-04 07:21:48","extension":"png","order_by":41,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":58871,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage10.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/f909413eb6b41084805854d5.png"},{"id":97666345,"identity":"e2fc415b-ec1c-46da-be58-7ca07be293ac","added_by":"auto","created_at":"2025-12-08 09:21:02","extension":"png","order_by":42,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":59494,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage11.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/6a3d8d025fe7c36d6b8ac42b.png"},{"id":97666787,"identity":"06dde350-dd65-462f-ac1a-6e045082175f","added_by":"auto","created_at":"2025-12-08 09:22:09","extension":"png","order_by":43,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":31759,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage12.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/5682b0d4f6fbc9e31babaf4b.png"},{"id":97666447,"identity":"9b20490b-9183-4569-800e-b03ce096e538","added_by":"auto","created_at":"2025-12-08 09:21:15","extension":"png","order_by":44,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":18322,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage15.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/28125e4aab15cb1c3ee8b49c.png"},{"id":97666985,"identity":"bdb5a416-5929-432c-b7ed-8b11b87563eb","added_by":"auto","created_at":"2025-12-08 09:22:34","extension":"png","order_by":45,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":20811,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage16.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/8ed7f4f39c2279f81cba3f63.png"},{"id":97667508,"identity":"c9a9bca7-7ea4-42ed-8f95-594fd4c600fd","added_by":"auto","created_at":"2025-12-08 09:23:41","extension":"png","order_by":46,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":9614,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage17.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/3226bae12aa0932e8e2ca617.png"},{"id":97666532,"identity":"fdc4fc2f-9ee2-4721-a274-0dc328867e1f","added_by":"auto","created_at":"2025-12-08 09:21:28","extension":"png","order_by":47,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":33432,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage18.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/6de493d64ece31be462348a3.png"},{"id":97667369,"identity":"d1c3af5a-ede5-4790-94f1-e86c76e3ae38","added_by":"auto","created_at":"2025-12-08 09:23:19","extension":"png","order_by":48,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":22122,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage19.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/9ea85f7b32bd838dc469db89.png"},{"id":97417109,"identity":"2c73adbd-e8d7-4542-b53c-a8e15c8e30df","added_by":"auto","created_at":"2025-12-04 07:21:48","extension":"png","order_by":49,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":43135,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage20.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/57a5c737459e20ca9e834bb2.png"},{"id":97417117,"identity":"b63848e9-137c-4cd1-98c8-4fe259e2f494","added_by":"auto","created_at":"2025-12-04 07:21:48","extension":"png","order_by":50,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":37520,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage21.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/762ead39eb60a71a625abd6f.png"},{"id":97417110,"identity":"b0beabba-312c-4ea2-ae70-9f411fc55485","added_by":"auto","created_at":"2025-12-04 07:21:48","extension":"xml","order_by":51,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":187036,"visible":true,"origin":"","legend":"","description":"","filename":"2d242c96728f444e9e03fed4b338f07b1structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/88b9f5d0d49f3ed5d45d38ff.xml"},{"id":97666518,"identity":"c160eb48-f5bc-4c99-bef1-fa049652f172","added_by":"auto","created_at":"2025-12-08 09:21:25","extension":"html","order_by":52,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":195235,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/b994e4fb252c28e335ce6f7d.html"},{"id":97417078,"identity":"6be8ff45-86bc-468c-b1df-e7e815cc1877","added_by":"auto","created_at":"2025-12-04 07:21:47","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":264863,"visible":true,"origin":"","legend":"\u003cp\u003eHierarchical Clustering Dendrogram\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/a6fe1f22d42f21db7b0a5490.png"},{"id":97667174,"identity":"78cf3eca-12ac-46fc-bf70-c214a85fb2f4","added_by":"auto","created_at":"2025-12-08 09:22:58","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":207485,"visible":true,"origin":"","legend":"\u003cp\u003eSimilarity Matrix\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/b5ddae93e4cdef38c47389a5.png"},{"id":97417072,"identity":"04d2d8cc-cd15-434e-bc73-1dbf1fdc4f70","added_by":"auto","created_at":"2025-12-04 07:21:47","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":81341,"visible":true,"origin":"","legend":"\u003cp\u003eIntertopic Distance Map\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/ff6ccfc5cd69349646967f2c.png"},{"id":97417075,"identity":"40ca7602-641d-46bc-abef-c0f6774d1b1b","added_by":"auto","created_at":"2025-12-04 07:21:47","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":102409,"visible":true,"origin":"","legend":"\u003cp\u003eTopics over Time\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/b9f5f444192b288cd2a7335a.png"},{"id":97417073,"identity":"9e7b04f6-711e-40d7-99b7-ae43a0ff2a29","added_by":"auto","created_at":"2025-12-04 07:21:47","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":32474,"visible":true,"origin":"","legend":"\u003cp\u003eDanmaku Interaction Patterns\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/7e5539abce79205784c1908e.png"},{"id":97666902,"identity":"4c9b1fe6-1982-47b3-a7ba-7352798ee3a0","added_by":"auto","created_at":"2025-12-08 09:22:23","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":116542,"visible":true,"origin":"","legend":"\u003cp\u003eAudio Feature Profiles\u003c/p\u003e","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/8956b8e6c8049327420cf5b9.png"},{"id":97666570,"identity":"43544bf9-ec47-4ce9-9269-40d670b11b78","added_by":"auto","created_at":"2025-12-08 09:21:36","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":107279,"visible":true,"origin":"","legend":"\u003cp\u003eVisual Feature Analysis\u003c/p\u003e","description":"","filename":"floatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/3c09ad85189d6ecfe2671261.png"},{"id":97417079,"identity":"ce4056a1-8b41-486e-ba20-70dd51b0f9da","added_by":"auto","created_at":"2025-12-04 07:21:47","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":264863,"visible":true,"origin":"","legend":"\u003cp\u003eHierarchical Clustering Dendrogram\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/896e42858ae144fe6a4c7c23.png"},{"id":97417080,"identity":"7029fa8e-1335-40e9-ad44-ca00790c7969","added_by":"auto","created_at":"2025-12-04 07:21:47","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":188709,"visible":true,"origin":"","legend":"\u003cp\u003eRegression Diagnostic Plots\u003c/p\u003e","description":"","filename":"floatimage8.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/e07c04571d8ca7836d18bd7b.png"},{"id":97417089,"identity":"6c792416-b11a-4b3c-b327-c14bef5bbe8b","added_by":"auto","created_at":"2025-12-04 07:21:47","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":207485,"visible":true,"origin":"","legend":"\u003cp\u003eSimilarity Matrix\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/5953e623322edefe0bf952b7.png"},{"id":97667220,"identity":"f08cf92c-62da-460f-87bd-cdee782ce75e","added_by":"auto","created_at":"2025-12-08 09:23:04","extension":"png","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":81341,"visible":true,"origin":"","legend":"\u003cp\u003eIntertopic Distance Map\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/894c0abe747b93225e2cf56d.png"},{"id":97666867,"identity":"7afd9185-fa92-4228-935e-ab61882859c7","added_by":"auto","created_at":"2025-12-08 09:22:17","extension":"png","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":193908,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eModel Complexity vs Performance\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"floatimage9.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/923df639641b2342edc74098.png"},{"id":97666830,"identity":"4b9358dd-40ec-4929-bd6c-7e49bfca7db1","added_by":"auto","created_at":"2025-12-08 09:22:11","extension":"jpeg","order_by":10,"title":"Figure 10","display":"","copyAsset":false,"role":"figure","size":239610,"visible":true,"origin":"","legend":"\u003cp\u003eSHAP Global Feature Importance - top 20 features\u003c/p\u003e","description":"","filename":"floatimage10.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/2c34d5a79669151441ec92cc.jpeg"},{"id":97666363,"identity":"cc62e26f-42fe-46b3-b67c-3b5f5dc6be18","added_by":"auto","created_at":"2025-12-08 09:21:04","extension":"png","order_by":10,"title":"Figure 10","display":"","copyAsset":false,"role":"figure","size":102409,"visible":true,"origin":"","legend":"\u003cp\u003eTopics over Time\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/bdbdf9e7415dd87bd0707d3d.png"},{"id":97417092,"identity":"61b0d93e-2cba-4765-80a4-288b01ac5e4a","added_by":"auto","created_at":"2025-12-04 07:21:47","extension":"png","order_by":11,"title":"Figure 11","display":"","copyAsset":false,"role":"figure","size":32474,"visible":true,"origin":"","legend":"\u003cp\u003eDanmaku Interaction Patterns\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/30a02f6b90e4832daa7c8f92.png"},{"id":97417096,"identity":"fe6da9d5-3589-43bc-b0f5-846cb49c0582","added_by":"auto","created_at":"2025-12-04 07:21:47","extension":"png","order_by":11,"title":"Figure 11","display":"","copyAsset":false,"role":"figure","size":294772,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSHAP Feature Dependence Plots\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"floatimage11.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/49b85834fd20f71bff875da6.png"},{"id":97417083,"identity":"e749ad04-73af-4fa8-889a-0807faa2908d","added_by":"auto","created_at":"2025-12-04 07:21:47","extension":"png","order_by":12,"title":"Figure 12","display":"","copyAsset":false,"role":"figure","size":116542,"visible":true,"origin":"","legend":"\u003cp\u003eAudio Feature Profiles\u003c/p\u003e","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/e2fb09ee3e78e32815570eac.png"},{"id":97417099,"identity":"02106982-fc42-44e5-9317-f6ea9b0b95f3","added_by":"auto","created_at":"2025-12-04 07:21:47","extension":"png","order_by":12,"title":"Figure 12","display":"","copyAsset":false,"role":"figure","size":189319,"visible":true,"origin":"","legend":"\u003cp\u003eFinal Model Performance\u003c/p\u003e","description":"","filename":"floatimage12.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/af42919942ded1810d613ebd.png"},{"id":97417081,"identity":"ba7b5a90-bec9-46c0-b28c-f5781e77fe92","added_by":"auto","created_at":"2025-12-04 07:21:47","extension":"png","order_by":13,"title":"Figure 13","display":"","copyAsset":false,"role":"figure","size":107279,"visible":true,"origin":"","legend":"\u003cp\u003eVisual Feature Analysis\u003c/p\u003e","description":"","filename":"floatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/2b3aa3adaa0f9174bb8cffcc.png"},{"id":97417086,"identity":"9b62e19d-ed77-4932-891e-56595a2ad60d","added_by":"auto","created_at":"2025-12-04 07:21:47","extension":"png","order_by":15,"title":"Figure 15","display":"","copyAsset":false,"role":"figure","size":188709,"visible":true,"origin":"","legend":"\u003cp\u003eRegression Diagnostic Plots\u003c/p\u003e","description":"","filename":"floatimage8.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/4680d4f89383ba34b2759047.png"},{"id":97666873,"identity":"1d375e8d-e2c4-4270-ae1c-cfdd19a3b857","added_by":"auto","created_at":"2025-12-08 09:22:17","extension":"png","order_by":16,"title":"Figure 16","display":"","copyAsset":false,"role":"figure","size":193908,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eModel Complexity vs Performance\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"floatimage9.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/20c4206d4725b28e3f4396f6.png"},{"id":97417108,"identity":"ffc3c05a-41fc-4a80-ae6e-a8433cd6b4fe","added_by":"auto","created_at":"2025-12-04 07:21:48","extension":"jpeg","order_by":17,"title":"Figure 17","display":"","copyAsset":false,"role":"figure","size":239610,"visible":true,"origin":"","legend":"\u003cp\u003eSHAP Global Feature Importance - top 20 features\u003c/p\u003e","description":"","filename":"floatimage10.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/c141452c8e8c45f4f4069ca7.jpeg"},{"id":97666578,"identity":"6b6dd3e4-3e8e-4467-92aa-857af979a8e2","added_by":"auto","created_at":"2025-12-08 09:21:37","extension":"png","order_by":18,"title":"Figure 18","display":"","copyAsset":false,"role":"figure","size":294772,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSHAP Feature Dependence Plots\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"floatimage11.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/a93f55ef2ce48931d4e01425.png"},{"id":97666480,"identity":"1c3bddec-bddc-4d07-82a4-78f6a88799f7","added_by":"auto","created_at":"2025-12-08 09:21:19","extension":"png","order_by":19,"title":"Figure 19","display":"","copyAsset":false,"role":"figure","size":189319,"visible":true,"origin":"","legend":"\u003cp\u003eFinal Model Performance\u003c/p\u003e","description":"","filename":"floatimage12.png","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/e5937ece32e04cd44cad9ff2.png"},{"id":97677549,"identity":"5ddcde5b-09a5-448f-99de-dfdc53a0662b","added_by":"auto","created_at":"2025-12-08 09:53:28","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":4890158,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8141450/v1/621a0e6f-eb73-48a8-91d0-9dc939922cd9.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"From Watching to Wishing: A Multimodal Computational Analysis of How PUGC Video-Danmaku Ecology Shapes Travel Intention","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eThe contemporary tourism landscape is undergoing a fundamental transformation driven by the global shift toward an experience economy where consumers increasingly prioritize authentic, personalized, and transformative travel experiences over standardized tourism products(B. Joseph Pine II and Gilmore James H. 1998; Chen 2025). This evolution has profoundly disrupted traditional marketing approaches, rendering top-down promotional strategies progressively ineffective due to their perceived lack of authenticity and unidirectional communication structure(Fotis et al. 2012; Leung et al. 2013). In response to these shifting dynamics, Professional User-Generated Content (PUGC) has emerged as a dominant force in shaping travel decisions, particularly among younger, digitally native tourists who seek peer validation and authentic experiences over commercial messaging(Pascual-Fraile et al. 2025)\u003c/p\u003e\u003cp\u003ePUGC represents a unique hybrid form that strategically combines the production values of professional content with the perceived authenticity and relatability of user-generated materials, creating a powerful influence mechanism that operates through parasocial relationships between creators and viewers(Huang 2024). On platforms like Bilibili and YouTube, these content creators have become critical nodes in the tourism information ecosystem, cultivating trust and influencing millions through carefully crafted multimodal narratives that blur the boundaries between entertainment and marketing(Horton and Richard Wohl 1956; Kim and Ko 2012). Despite the acknowledged importance of PUGC in contemporary tourism marketing, a significant gap persists in our empirical understanding of the specific mechanisms through which this influence operates, with the pathway from passive viewing to concrete travel intention formation remaining largely unexplored(Li and Tu 2024a).\u003c/p\u003e\u003cp\u003eThis empirical gap becomes particularly pronounced when considering the unique media ecology of platforms like Bilibili, which feature Danmaku, a distinctive form of real-time, synchronous commentary that appears as scrolling text overlays directly on the video screen, transforming solitary viewing into a collective social experience (Yang 2020a; Cauchard et al. 2024). Unlike traditional asynchronous comments that appear below or beside video content, Danmaku creates what scholars term \"pseudo-synchronous co-viewing,\" where viewers experience the illusion of watching together with thousands of others, their reactions and interpretations becoming part of the content itself (Wu et al. 2019; Liu 2024). This technological affordance fundamentally alters the information processing environment in which travel decisions form, introducing continuous social validation signals that may amplify, moderate, or potentially undermine the primary content's influence(Yang et al. 2022).\u003c/p\u003e\u003cp\u003eThe role of Danmaku in high-involvement decision-making processes like travel planning remains entirely unexplored in the academic literature, despite strong theoretical reasons to expect different dynamics than those observed in low-stakes entertainment or e-commerce contexts where most Danmaku research has focused (Lyu 2021; Fan et al. 2023). Travel decisions involve substantial financial commitments, extended time horizons, and significant personal risks, suggesting that the casual social proof mechanisms effective for impulse purchases may operate differently when viewers contemplate international travel or significant tourism investments. Furthermore, the multimodal complexity of tourism videos, which combines scenic visuals, ambient sounds, narrative voiceovers, and cultural content, creates a rich but potentially overwhelming information environment where Danmaku might serve as either helpful interpretation guides or distracting noise that impedes decision-making.\u003c/p\u003e\u003cp\u003eAddressing these critical gaps requires methodological innovation beyond traditional approaches. Survey-based studies and manual content analysis, while valuable for capturing subjective experiences and semantic meanings, cannot adequately capture the dynamic, multimodal, and temporally synchronized nature of PUGC-Danmaku ecology. Recent advances in computational methods, particularly in computer vision, natural language processing, and multimodal machine learning, offer unprecedented opportunities to systematically investigate these complex phenomena at scale while maintaining analytical rigor(Baltrusaitis et al. 2019; Xu et al. 2023). By extracting and analyzing hundreds of features across visual, auditory, and textual modalities, we can move beyond speculation about influence mechanisms to empirical measurement of how specific multimodal configurations and interaction patterns shape viewer responses.\u003c/p\u003e\u003cp\u003eThis study therefore aims to open the \"black box\" of PUGC influence by implementing an innovative multimodal computational framework that systematically investigates the complex interplay between video content and real-time audience interaction in shaping travel intentions. Through the analysis of 2,650 tourism videos from Bilibili containing millions of Danmaku comments, we seek to answer three fundamental research questions that address both theoretical and practical concerns in digital tourism marketing. First, we investigate how the multimodal profiles of different types of PUGC travel videos systematically differ, moving beyond intuitive categorizations to data-driven discovery of content strategies. Second, we examine how these distinct multimodal profiles influence viewers' travel intentions, testing whether traditional linear frameworks adequately capture these relationships or whether more complex mechanisms operate. Third, we assess the specific role and relative importance of the Danmaku interaction ecology in this influence process, determining whether synchronized social commentary represents mere digital noise or a critical influence pathway that tourism marketers must understand and leverage.\u003c/p\u003e"},{"header":"2. Literature Review","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003e2.1 The Evolution of Tourism Influence in Digital Contexts\u003c/h2\u003e\u003cp\u003eThe transformation of tourism marketing from traditional broadcast approaches to participatory digital narratives represents more than a simple channel shift; it fundamentally alters the mechanisms through which destinations build awareness, shape perceptions, and ultimately influence travel decisions(Pop et al. 2022). Professional User-Generated Content has emerged at the intersection of this transformation, combining professional production capabilities with the authenticity markers (Chen 1986; Luo 1986) that contemporary consumers use to assess credibility in an environment saturated with commercial messaging. While tourism scholars have extensively documented the rise of influencer marketing and its impact on destination image formation (Gretzel 2018), critical gaps remain in our understanding of how PUGC's distinctive affordances, particularly its multimodal orchestration and interactive overlays, mechanistically translate viewing experiences into concrete travel intentions.\u003c/p\u003e\u003cp\u003eRecent empirical investigations provide compelling but incomplete evidence for PUGC's effectiveness in tourism contexts. Li (Li and Tu 2024a)and Tu (2024) offer the most direct comparative evidence through experimental manipulation, demonstrating that PUGC's advantage over both amateur user-generated content and professional advertising depends critically on destination type and viewer characteristics. For hedonic destinations emphasizing pleasure and experience, authenticity signals emerged as the dominant influence pathway, while utilitarian destinations focusing on practical benefits gained more from professional production quality. This contingency challenges universal claims about influencer superiority and suggests that the frequently cited \"authenticity-professionalism blend\" (Peinado and Shim 2024) operates through more nuanced mechanisms than previously theorized. Their finding that credibility and usefulness mediate these effects (Li and Tu 2024b) aligns with dual-processing frameworks from consumer psychology(Samson and Voyer 2012; Wyer and Kardes 2020), yet leaves unexamined the specific audiovisual techniques through which creators achieve this delicate balance between professional polish and authentic expression.\u003c/p\u003e\u003cp\u003eThe psychological mechanisms underlying PUGC influence extend beyond surface-level credibility assessments to encompass parasocial relationships: the one-sided emotional bonds that viewers develop with media personalities through repeated exposure and perceived intimacy(Horton and Richard Wohl 1956). Contemporary research confirms that these parasocial bonds systematically condition how viewers process both informational and affective content, creating a relationship context that transforms objective destination information into personally relevant travel inspiration (Silaban et al. 2023; Wu et al. 2024). Notably, Nguyen et al. (2024)(Nguyen et al. 2024) provide particularly relevant evidence by demonstrating that inspiration fully mediates the relationship between professionalism and travel planning while only partially mediating sincerity effects, suggesting heterogeneous pathways from creator characteristics to behavioral outcomes. This differentiation becomes crucial for understanding PUGC's influence in high-involvement decisions like travel, where emotional resonance may override rational evaluation of destination attributes(Pham 2007), particularly when viewers have developed strong parasocial bonds with creators(Gong and Li 2017; Liu et al. 2024).\u003c/p\u003e\u003cp\u003eHowever, a critical limitation pervades the existing PUGC tourism literature (Chemin et al. 2025) that undermines theoretical development and practical application. Despite widespread recognition that video represents PUGC's primary medium (Polat et al. 2023) and theoretical claims that multimodal storytelling drives its effectiveness(Shah et al. 2018), empirical studies consistently treat \"video\" as an undifferentiated, monolithic format without examining the specific affordances that distinguish successful from unsuccessful content. Researchers typically measure exposure to \"travel vlogs\" or \"destination videos\" (Buckingham 2009; Ross et al. 2011)without operationalizing the production choices comprising the viewing experience: camera angles, editing pace, music selection, color grading, and narrative structure. This oversight is problematic as production quality shapes viewer perceptions in complex ways. The relationship between amateur and polished presentation styles remains underexplored, with (Ward et al. 2020) Ward et al. (2020) treating production style as a binary category rather than decomposing it into measurable multimodal components. This approach obscures which specific audiovisual elements enhance or undermine authenticity perceptions, limiting our understanding of how production choices influence tourism marketing effectiveness.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e\u003ch2\u003e2.2 Danmaku as a Unique Mechanism of Synchronous Social Influence\u003c/h2\u003e\u003cp\u003eThe phenomenon of Danmaku introduces an entirely new dimension to digital content consumption that fundamentally challenges traditional models of media influence developed for passive, individual viewing contexts. Originating from Japanese \"Niconico\" video culture and adapted with distinctive characteristics in Chinese platforms like Bilibili, Danmaku represents more than a technological feature; it constitutes a new form of collective viewing that transforms how audiences process and interpret media content (Fang et al. 2018). Unlike conventional commenting systems where feedback appears spatially separated from content and temporally divorced from specific moments, Danmaku overlays create what Wu et al. (2019)(Wu et al. 2019) term \"pseudo-synchronous co-viewing,\" where comments from different times appear simultaneously on screen, creating the psychological experience of watching alongside a crowd despite physical and temporal separation.\u003c/p\u003e\u003cp\u003eThe theoretical foundations for understanding Danmaku's influence on viewer cognition and behavior derive from two complementary psychological mechanisms that operate simultaneously during viewing experiences. Emotional contagion theory, first formalized by Hatfield et al. (1993)(Hatfield et al. 1993), predicts that exposure to others' emotional expressions triggers automatic mimicry and affective convergence, leading individuals to \"catch\" the emotions displayed by those around them. Controlled experimental studies confirm this prediction in Danmaku contexts, with Kamei et al. (2012)(KAMEI et al. 2012) demonstrating measurable affect transfer when viewers are exposed to emotionally congruent overlays, while Zhou et al. (2018)(Zhou et al. 2018) show that excitement-laden Danmaku significantly increases viewers' willingness to send virtual gifts to content creators. These findings suggest that the emotional tone of Danmaku may shape viewers' affective responses to tourism content independent of the primary video's emotional valence, potentially amplifying positive responses to destinations or neutralizing negative impressions through collective enthusiasm.\u003c/p\u003e\u003cp\u003eComplementing emotional contagion, social proof theory provides a cognitive pathway through which Danmaku influences decision-making by serving as heuristic cues about content value and appropriate responses(Cialdini 1993). When viewers observe numerous comments expressing desire to visit a destination or praising specific attractions, these expressions function as social validation that reduces uncertainty and legitimizes similar responses. Yuan et al. (2024)(Yuan et al. 2024) provide sophisticated time-varying evidence for this mechanism, showing that Danmaku density bursts correlate with increased engagement metrics even after controlling for content quality, suggesting that the mere presence of active commentary signals content worth attending to. Pan's (2023) (Pan 2023)qualitative analysis reveals that users explicitly interpret Danmaku as \"being with\" rather than merely observing others, transforming their viewing from isolated consumption into participatory experience where collective reactions guide individual interpretation.\u003c/p\u003e\u003cp\u003eHowever, emerging evidence suggests important boundary conditions that complicate straightforward application of these theories to tourism contexts((Brian) Lin et al. 2022). The vast majority of Danmaku research has examined low-stakes, hedonic consumption environments, such as gaming streams, variety shows, and e-commerce broadcasts (Chen et al. 2015; Pan 2023), where decisions involve minimal risk and immediate gratification. In these contexts, the \"lively atmosphere\" created by dense Danmaku reliably increases engagement, purchase intention, and platform loyalty (Fang et al. 2018; Yuan et al. 2024). Yet Zhang and Ruan (2024)(Rahwan et al. 2024) recently uncovered a critical counter-effect that challenges universal application of social proof principles: excessive comment homogeneity breeds psychological reactance, where viewers resist perceived social pressure and assert autonomy by rejecting popular opinions. This reactance effect may be particularly pronounced in high-involvement decisions like travel planning, where substantial financial and temporal investments make viewers more sensitive to manipulation attempts and less susceptible to momentary social influence.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec5\" class=\"Section2\"\u003e\u003ch2\u003e2.3 Methodological Limitations and the Computational Turn\u003c/h2\u003e\u003cp\u003eThe complexity inherent in PUGC-Danmaku ecology exposes fundamental limitations in tourism research's traditional methodological toolkit, revealing a troubling disconnect between theoretical sophistication and empirical capability. Contemporary theories emphasize multimodal orchestration, temporal dynamics, and real-time social interaction as defining features of digital tourism influence, yet the field's dominant methods, namely surveys, interviews, and manual content analysis, were designed for static, monomodal data and cannot adequately capture these phenomena (Naab et al. 2019). This methodological constraint has created a literature rich in conceptual frameworks but impoverished in mechanistic evidence, where researchers theorize about complex influence processes they cannot empirically observe or measure.\u003c/p\u003e\u003cp\u003eSurvey-based studies, while remaining the dominant approach in tourism research, suffer from systematic validity threats that become particularly acute when investigating digital media consumption(Luo et al. 2021; Otto et al. 2022). Naab et al. (2019) provide sobering evidence that retrospective self-reports of even simple behaviors like video viewing duration show average errors of 40\u0026ndash;60% when compared to server logs, with systematic biases toward overestimation of socially desirable behaviors and underestimation of passive consumption. For complex multimodal experiences involving rapid scene changes, background music, narrative voiceovers, and scrolling text overlays, the cognitive burden of accurate recall likely increases exponentially, rendering self-report data highly suspect(Pink and Newton 2020). More fundamentally, surveys cannot capture the temporal dynamics central to digital influence: the specific moments when visual revelations align with musical crescendos and collective commentary to create peak emotional experiences that crystallize travel intentions.\u003c/p\u003e\u003cp\u003eManual content analysis faces even more severe limitations when applied to PUGC-Danmaku ecology, as the sheer volume and complexity of data overwhelm human coding capacity(Grimmer and Stewart 2013; Alaei et al. 2019). While sophisticated frameworks like Schwenzow et al.'s (2021)(Schwenzow et al. 2021) multimodal coding scheme provide systematic protocols for analyzing video content across multiple dimensions, the authors acknowledge that labor requirements increase \"immensely\" with video length and modal complexity(Peinado et al. 2020), rendering them impractical for the thousands of hours typical in PUGC corpora. The challenge compounds when considering Danmaku, where a single popular video may contain hundreds of thousands of time-coded comments that would require years of manual coding to analyze comprehensively. Attempts to address scale through crowdsourcing sacrifice the interpretive nuance and contextual understanding that justify human coding over automated methods(Larson et al. 2014), while inter-coder reliability deteriorates rapidly as the number of variables and modalities increases(Bayerl and Paul 2011).\u003c/p\u003e\u003cp\u003eThe emergence of computational methods, particularly advances in deep learning and multimodal machine learning, offers a principled solution to these methodological constraints while enabling investigation of previously unobservable phenomena. Contemporary computer vision models can extract hundreds of visual features from video frames, including object detection, scene classification, aesthetic attributes, and emotional expressions, with superhuman accuracy and perfect reliability (He et al. 2016). Natural language processing techniques can analyze millions of comments to identify semantic patterns, emotional valences, and social network structures that would be impossible to detect manually(Vaswani et al. 2023). Most critically, multimodal fusion architectures can model the complex interactions between visual, auditory, and textual streams, capturing cross-modal dependencies and temporal dynamics that linear statistical methods cannot represent(Baltrusaitis et al. 2019; Xu et al. 2023) .\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\u003ch2\u003e2.4 An Integrated Theoretical Framework: The Computationally Extended S-O-R Model\u003c/h2\u003e\u003cp\u003eBuilding on the foundational Stimulus-Organism-Response paradigm that has guided environmental psychology and consumer behavior research for decades (Mehrabian and Russell 1980), we propose a Computationally Extended S-O-R Model specifically designed to capture the unique dynamics of PUGC-Danmaku influence in tourism contexts. This framework addresses critical limitations in existing S-O-R applications to digital tourism, which typically treat stimuli as static and independent(Asyraff et al. 2023), assume linear additive effects, and rely on self-reported organism states that suffer from severe measurement error(Li et al. 2015; Y\u0026uuml;ksel 2017). Our extension integrates insights from cognitive load theory(Sweller 1988), social proof theory, and parasocial relationship theory(Horton and Richard Wohl 1956) while enabling empirical validation through computational measurement of previously unobservable constructs.\u003c/p\u003e\u003cp\u003eThe stimulus layer in our framework recognizes that viewers encounter not a single coherent message but multiple competing information streams that vie for limited cognitive resources(Turner and O\u0026rsquo;Leary 2012). Primary video content delivers multimodal stimuli through visual semantics captured by CLIP embeddings(Radford et al. 2003), acoustic features including spectral characteristics and temporal patterns, and narrative structures reflected in speech transcription and pacing. Simultaneously, Danmaku overlays provide continuous social stimuli through comment density patterns, semantic content, emotional valence, and temporal clustering that may reinforce, contradict, or reframe the primary content's message. Unlike traditional S-O-R applications that assume stimuli combine additively(Thein et al. 2022), our framework explicitly models modal competition through cross-modal correlation analysis and temporal alignment metrics, recognizing that simultaneous information streams may interfere with rather than enhance each other.\u003c/p\u003e\u003cp\u003eThe organism layer represents the critical mediating processes through which external stimuli translate into behavioral responses, but rather than relying on error-prone self-reports(Podsakoff and Organ 1986; Stone and Shiffman 2002), we infer cognitive and affective states from objective behavioral traces. Cognitive load, traditionally measured through subjective ratings(Skulmowski and Rey 2015) or secondary task performance(Haji et al. 2015; Ehlers 2020), emerges from the computational complexity of multimodal features(Brunken et al. 2003; Ross et al. 2003), wherein the variance in visual semantics indicates intrinsic load from content difficulty, cross-modal asynchrony reflects extraneous load from poor design(Kan 2023), and the semantic coherence between video and Danmaku suggests germane load from meaningful elaboration(Leng et al. 2016). Social processing manifests through Danmaku clustering patterns that reveal when viewers collectively attend to specific moments, sentiment cascades that show emotional contagion(Yu et al. 2025), and semantic convergence that indicates shared interpretation. Parasocial bonding, while not directly observable, leaves traces in the personalization of comments, frequency of creator references, and persistence of viewing across a creator's catalog(Fazli-Salehi et al. 2022; Tan-intaraarj 2024).\u003c/p\u003e\u003cp\u003eThe response layer captures behavioral outcomes through platform-native indicators rather than artificial research instruments, enhancing ecological validity(SU et al. 2021) while enabling large-scale measurement(Diehl et al. 2017). Travel intention, our primary outcome, emerges from computational analysis of viewer comments(Mehra 2023) where expressions like \"I want to visit\" and \"added to my bucket list\" represent natural behavioral signals uncontaminated by research demand effects(Caulley 1994; Bailenson et al. 2004). Engagement patterns including viewing duration, replay behavior, and sharing actions provide complementary indicators of content influence(Wu et al. 2018; Anh 2024), while the temporal distribution of intention expressions reveals which specific moments trigger decision crystallization(Li et al. 2025). This behavioral approach acknowledges that expressed interest may not equal booking behavior but captures the crucial early stages of travel decision-making where awareness transforms into consideration(Dimitriou and AbouElgheit 2019).\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e\u003ch2\u003e2.5 Research Hypotheses Development\u003c/h2\u003e\u003cp\u003eThe theoretical foundations and methodological innovations discussed above converge on fundamental questions about how PUGC videos and their Danmaku overlays jointly shape travel intentions through multimodal orchestration and social dynamics. Our hypothesis development proceeds from three interconnected theoretical gaps that emerged from the literature review. First, despite recognition that PUGC creators employ different communication strategies(Pertiwi and Sanusi 2023), no empirical evidence exists for whether these strategies manifest as measurably distinct multimodal profiles or represent random variation(Zhu et al. 2025). Second, while theory predicts that multimodal features and cognitive processing should influence travel decisions(Jun and Vogt 2013), the functional form of these relationships, whether linear, threshold-based, or interactive, remains unknown. Third, although Danmaku theoretically provides social influence, its relative importance compared to primary content features in high-involvement decisions like travel has never been tested(Filieri et al. 2025).\u003c/p\u003e\u003cp\u003eBuilding on cognitive load theory's prediction that information complexity systematically affects processing and retention (Sweller 1988), combined with evidence that different tourism content serves distinct communication goals, we propose that PUGC videos will cluster into categories with characteristic multimodal signatures. These categories should differ not only in surface features like topic and style but in fundamental properties including visual dynamism, audio characteristics, narrative structure, and crucially, the cognitive demands they impose on viewers. Furthermore, these differences should manifest in how audiences engage through Danmaku, with some content types triggering immediate reactive commentary while others promote reflective discussion. Therefore:\u003c/p\u003e\u003cp\u003eH1: Data-driven categories of PUGC travel videos will exhibit systematically different multimodal and cognitive profiles that reflect distinct communication strategies.\u003c/p\u003e\u003cp\u003eH1a: Categories will show significant differences in objective multimodal features including visual dynamism, audio characteristics, and Danmaku density patterns, demonstrating that content types employ distinct production approaches rather than random variation.\u003c/p\u003e\u003cp\u003eH1b: Categories will impose different levels of cognitive load as computationally derived from multimodal feature complexity and temporal variation, indicating that cognitive demands represent a strategic choice rather than incidental byproduct.\u003c/p\u003e\u003cp\u003eThe relationship between multimodal features and travel intention likely operates through complex mechanisms that traditional linear models cannot capture. Cognitive load theory suggests threshold effects where information must reach sufficient complexity to engage deep processing but not so much as to overwhelm capacity. Social proof theory predicts that Danmaku influence depends on reaching critical mass where collective enthusiasm becomes self-reinforcing. Parasocial relationship theory implies that influence accumulates through repeated exposure to consistent creator personas rather than linearly with each video. These theoretical predictions, combined with evidence of non-linear effects in related domains, suggest:\u003c/p\u003e\u003cp\u003eH2: The multimodal and cognitive profiles of PUGC videos will significantly predict viewers' travel intentions through identifiable causal pathways.\u003c/p\u003e\u003cp\u003eH2a: Multimodal feature dimensions and cognitive load metrics will demonstrate significant associations with travel intention expressions, though these relationships may be non-linear and involve threshold effects invisible to traditional regression approaches.\u003c/p\u003e\u003cp\u003eH2b: Temporal causality analysis will reveal systematic lead-lag relationships between content features and audience reactions, with visual elements triggering audio responses that subsequently generate Danmaku commentary, providing evidence for orchestrated influence sequences.\u003c/p\u003e\u003cp\u003eThe integration of multiple information streams in PUGC-Danmaku ecology suggests that influence emerges from the gestalt rather than individual components. Multimodal learning theory demonstrates that cross-modal integration often yields superior outcomes compared to single modalities when information streams are complementary rather than redundant. In tourism contexts, visual beauty might capture attention, narrative provides meaning, music sets emotional tone, and Danmaku offers social validation, with each modality being insufficient alone but powerful in combination. Moreover, the unique cognitive demands of processing simultaneous video and text streams may create distinctive influence patterns where cognitive engagement becomes the bottleneck determining whether multimodal richness translates into behavioral influence. Therefore:\u003c/p\u003e\u003cp\u003eH3: Multimodal integration will provide superior prediction of travel intention compared to unimodal approaches, with interactive features and cognitive processing playing critical roles.\u003c/p\u003e\u003cp\u003eH3a: Machine learning models using integrated multimodal features will significantly outperform single-modality baselines in predicting travel intention expressions, demonstrating that influence emerges from cross-modal interactions rather than additive effects.\u003c/p\u003e\u003cp\u003eH3b: Feature importance analysis will reveal that audience interaction metrics and cognitive load measures are among the most influential predictors, potentially surpassing traditional content quality indicators like visual aesthetics or production values.\u003c/p\u003e\u003c/div\u003e"},{"header":"3. Methodology","content":"\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e\u003ch2\u003e3.1 Data Collection and Computational Framework\u003c/h2\u003e\u003cp\u003eThis investigation employed a comprehensive computational framework to analyze PUGC tourism videos and their associated Danmaku commentary from Bilibili, China's leading video-sharing platform characterized by its distinctive synchronous commenting system. The selection of Bilibili as our data source was motivated by three factors that make it ideal for investigating multimodal influence mechanisms: the platform's native integration of Danmaku creates a natural laboratory for studying synchronized social viewing, its predominantly young user base (82% aged 18\u0026ndash;35) represents the digitally native tourists driving industry transformation, and its open API enables systematic data collection at scales impossible with manual methods. The data collection proceeded through multiple phases between January and February 2025, beginning with exploratory sampling to understand content diversity and concluding with targeted collection to ensure adequate representation across emergent categories.\u003c/p\u003e\u003cp\u003eThis study employed a comprehensive computational framework to analyze PUGC tourism videos from Bilibili, China's leading video-sharing platform characterized by its distinctive Danmaku commenting system. The platform's unique ecology, where time-synchronized comments overlay video content, provides an ideal natural laboratory for investigating how multimodal content and real-time social interactions jointly influence travel decisions.\u003c/p\u003e\u003cp\u003eData collection proceeded through two complementary phases between January and February 2025. The initial exploratory phase cast a wide net using broad tourism-related Chinese keywords (\"旅游\" [tourism], \"旅行\" [travel], \"游玩\" [leisure travel]) to understand the landscape of tourism content, yielding metadata for 1,640 videos. After quality screening removed duplicates and videos shorter than 60 seconds, 1,400 videos remained for analysis.\u003c/p\u003e\u003cp\u003eTo avoid imposing researcher-defined categories while ensuring systematic organization, we employed unsupervised clustering via BERTopic(Grootendorst 2020). Video titles and descriptions underwent segmentation using jieba with a custom tourism vocabulary, then were embedded into 1,024-dimensional semantic space using Conan-embedding-v1. The clustering pipeline utilized UMAP for dimensionality reduction (n_neighbors\u0026thinsp;=\u0026thinsp;15, n_components\u0026thinsp;=\u0026thinsp;10, min_dist\u0026thinsp;=\u0026thinsp;0.1) to preserve both local and global structure, followed by HDBSCAN (min_cluster_size\u0026thinsp;=\u0026thinsp;30, min_samples\u0026thinsp;=\u0026thinsp;10) for density-based cluster identification. This process revealed 67 fine-grained topics that were hierarchically consolidated into three macro-categories using Ward's linkage, representing distinct communication strategies in PUGC tourism content.\u003c/p\u003e\u003cp\u003eThe emergent categorization revealed imbalanced representation across categories, motivating a targeted second collection phase using category-specific keywords. This yielded a final corpus of 2,650 videos: 991 Destination Attraction videos focusing on scenic locations, 1,116 Tourism Narrative Communication videos emphasizing personal experiences, and 543 Overseas Cultural Experience videos highlighting international encounters. For each video, we collected complete video files, all Danmaku comments with precise timestamps (N\u0026thinsp;=\u0026thinsp;475,253,138), and hierarchical viewer comments (N\u0026thinsp;=\u0026thinsp;11,849,985), creating a rich multimodal dataset for analysis.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e\u003ch2\u003e3.2 Multimodal Feature Extraction and Engineering\u003c/h2\u003e\u003cp\u003eFeature extraction required balancing temporal granularity with computational efficiency while preserving content dynamics. We implemented a sliding window approach with 5-second windows and 2.5-second overlap, a configuration that captures content transitions while maintaining sufficient stability for reliable feature extraction. This approach generated overlapping segments that preserve temporal continuity, essential for understanding how multimodal patterns evolve throughout videos.\u003c/p\u003e\u003cp\u003eVisual feature extraction leveraged three complementary approaches to capture semantic and aesthetic properties. CLIP ViT-B/32 generated 512-dimensional semantic embeddings for each frame, enabling quantification of visual diversity and thematic coherence. ResNet50 pre-trained on Places365 identified environmental contexts from 365 scene categories particularly relevant to tourism. YOLOv8n provided object detection for measuring visual complexity through entity counts. Aggregation from frame-level to window-level employed both statistical summarization (mean, standard deviation, range) and temporal modeling (trend coefficients, change rates), yielding 627 visual feature dimensions that comprehensively characterize visual content.\u003c/p\u003e\u003cp\u003eAudio analysis extracted 75 acoustic features via librosa, encompassing multiple perceptual dimensions. These included 13 MFCCs with first and second derivatives for timbral characterization essential for speech-music discrimination, spectral features (centroid, bandwidth, contrast, rolloff, flatness) differentiating ambient soundscapes from narration, temporal features (zero-crossing rate, tempo, onset patterns) indicating activity levels, and tonal features (chroma vectors, harmonic-percussive separation) distinguishing musical accompaniment from environmental sounds. This comprehensive acoustic profile captures production quality, emotional tone, and information density independent of semantic content.\u003c/p\u003e\u003cp\u003eDanmaku processing addressed the unique challenges of time-synchronized, overlapping text streams. Within each 5-second window, we aggregated all comments while preserving their temporal context. Text preprocessing employed jieba segmentation optimized for informal online language, supplemented by custom dictionaries for tourism terminology and internet slang. Each window's aggregated Danmaku was characterized through 1,024-dimensional embeddings (Conan-embedding-v1), density metrics (raw count, unique users, temporal clustering coefficients), and linguistic features (sentiment scores, lexical diversity, semantic coherence with video content). These features capture both quantitative and qualitative aspects of collective viewing experiences.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\u003ch2\u003e3.3 Multimodal Fusion and Cognitive Load Operationalization\u003c/h2\u003e\u003cp\u003eThe integration of heterogeneous modalities required sophisticated fusion techniques that preserve modality-specific information while capturing cross-modal interactions. Our two-stage fusion architecture addresses the fundamental challenge of comparing features from different measurement spaces while modeling their complex interdependencies.\u003c/p\u003e\u003cp\u003eStage one employed Deep Canonical Correlation Analysis (DCCA) to learn aligned representations between visual and audio modalities. The architecture consisted of parallel neural networks processing visual features (627 dimensions) and audio features (75 dimensions), both utilizing ReLU activation and dropout regularization. Training optimized these networks to maximize correlation between their 256-dimensional output representations.\u003c/p\u003e\u003cp\u003eStage two integrated the aligned audiovisual representations with Danmaku embeddings using a Transformer architecture that captures complex inter-modal dependencies through self-attention mechanisms. With eight attention heads enabling simultaneous focus on different relationship aspects and three encoder layers providing hierarchical representation learning, this architecture produced 256-dimensional fused representations. Layer normalization and dropout (p\u0026thinsp;=\u0026thinsp;0.1) ensured stable training and generalization, yielding holistic multimodal representations for each temporal window.\u003c/p\u003e\u003cp\u003eBuilding on Cognitive Load Theory's three-component framework (Sweller 1988), we operationalized cognitive dimensions through computational metrics. Intrinsic load, representing content's inherent difficulty, was measured through three components:\u003c/p\u003e\u003cp\u003e\u003cb\u003eContent Complexity (CC)\u003c/b\u003e:\u003cdiv id=\"Equ1\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ1\" name=\"EquationSource\"\u003e\n$$\\:\\begin{array}{c}CC=\\frac{\\sigma\\:\\left({F}_{f}\\right)}{\\mu\\:\\left({F}_{f}\\right)}\\end{array}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e1\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003eSemantic Density (SD)\u003c/b\u003e:\u003cdiv id=\"Equ2\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ2\" name=\"EquationSource\"\u003e\n$$\\:\\begin{array}{c}SD=\\text{Var}\\left({F}_{t}\\right)\\end{array}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e2\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003eConceptual Difficulty (CD)\u003c/b\u003e:\u003cdiv id=\"Equ3\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ3\" name=\"EquationSource\"\u003e\n$$\\:\\begin{array}{c}CD=\\frac{\\text{l}\\text{o}\\text{g}(DI\\times\\:DD+1)}{10}\\end{array}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e3\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003ewhere \u003cem\u003eF\u003c/em\u003e\u003csub\u003ef\u003c/sub\u003e represents fused features, \u003cem\u003eF\u003c/em\u003e\u003csub\u003et\u003c/sub\u003e represents text-aligned features, σ denotes standard deviation, \u0026micro; denotes mean, Var denotes variance, \u003cem\u003eDI\u003c/em\u003e is demand intensity, and \u003cem\u003eDD\u003c/em\u003e is demand diversity.\u003c/p\u003e\u003cp\u003eExtraneous load, representing presentation-imposed difficulty, was captured through:\u003c/p\u003e\u003cp\u003e\u003cb\u003eModal Interference (MI)\u003c/b\u003e:\u003cdiv id=\"Equ4\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ4\" name=\"EquationSource\"\u003e\n$$\\:\\begin{array}{c}MI=1-\\left|\\rho\\:\\right({F}_{v},{F}_{a}\\left)\\right|\\end{array}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e4\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003ePresentation Complexity (PC)\u003c/b\u003e:\u003cdiv id=\"Equ5\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ5\" name=\"EquationSource\"\u003e\n$$\\:\\begin{array}{c}PC=\\sigma\\:\\left({F}_{v}\\right)\\end{array}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e5\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003ewhere \u003cem\u003eF\u003c/em\u003e\u003csub\u003ev\u003c/sub\u003e represents video-aligned features, \u003cem\u003eF\u003c/em\u003e\u003csub\u003ea\u003c/sub\u003e represents audio-aligned features, and ρ denotes Pearson correlation coefficient. These operationalizations transform theoretical constructs into measurable quantities suitable for empirical analysis.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e\u003ch2\u003e3.4 Travel Intention Measurement Through Computational Text Analysis\u003c/h2\u003e\u003cp\u003eMeasuring travel intention in natural viewing contexts required innovation beyond traditional survey approaches. We developed a computational method to extract intention signals from 11.8\u0026nbsp;million viewer comments, providing unprecedented scale while maintaining ecological validity since comments represent spontaneous expressions uninfluenced by research instruments.\u003c/p\u003e\u003cp\u003eOur dual-stage approach balanced coverage with precision. The rule-based stage employed 47 carefully developed patterns capturing Chinese travel intention expressions, including direct statements (\"想去\" [want to go], \"种草了\" [added to list]), planning language (\"打算去\" [planning to visit], \"安排上\" [scheduling it]), and destination inquiries (\"在哪里\" [where is this], \"怎么去\" [how to get there]). However, keyword matching alone produced substantial false positives from conditional statements, quotations, and sarcasm. Therefore, a second-stage RoBERTa model fine-tuned on manually annotated comments distinguished genuine intentions from other travel-related language use.\u003c/p\u003e\u003cp\u003eThe final intention score was calculated as:\u003c/p\u003e\u003cp\u003e\u003cb\u003eTravel Intention Score (TIS)\u003c/b\u003e:\u003cdiv id=\"Equ6\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ6\" name=\"EquationSource\"\u003e\n$$\\:\\begin{array}{c}TI{S}_{i}=\\frac{{V}_{i}}{{C}_{i}}\\end{array}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e6\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003ewhere \u003cem\u003eV\u003c/em\u003e\u003csub\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sub\u003e represents validated intention expressions for video \u003cem\u003ei\u003c/em\u003e and \u003cem\u003eC\u003c/em\u003e\u003csub\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sub\u003e \u003cem\u003e\u003c/em\u003erepresents total comments for video \u003cem\u003ei\u003c/em\u003e. This provides a continuous measure of the proportion of viewers moved to express travel interest.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec13\" class=\"Section2\"\u003e\u003ch2\u003e3.5 Statistical Analysis and Model Development\u003c/h2\u003e\u003cp\u003ePrior to analysis, we conducted comprehensive assumption diagnostics. Homogeneity of variances was assessed using Levene's test for subsequent ANOVA procedures. Regression diagnostics included Shapiro-Wilk tests for residual normality, residual versus fitted value plots for homoscedasticity evaluation, and Variance Inflation Factors for multicollinearity detection. When violations were detected, we employed appropriate robust methods or non-parametric alternatives.\u003c/p\u003e\u003cp\u003eTo examine differences in multimodal profiles across video categories, one-way ANOVA was conducted for each of the 102 extracted features, with video category as the independent variable. Effect sizes were calculated using eta-squared to assess practical significance, and post-hoc comparisons employed Tukey's HSD test. The Benjamini-Hochberg procedure controlled false discovery rate at 0.10 to balance Type I error control with statistical power.\u003c/p\u003e\u003cp\u003eTo investigate the relationship between multimodal features and travel intention, we specified a multiple linear regression model using selected key features:\u003cdiv id=\"Equ7\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ7\" name=\"EquationSource\"\u003e\n$$\\:\\begin{array}{c}{Y}_{i}={\\beta\\:}_{0}+{\\beta\\:}_{1}{F}_{i1}+{\\beta\\:}_{2}{F}_{i2}+{\\beta\\:}_{3}{C}_{i1}+{\\beta\\:}_{4}{C}_{i2}+{\\delta\\:}_{1}{D}_{i1}+{\\delta\\:}_{2}{D}_{i2}+{\\epsilon\\:}_{i}\\end{array}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e7\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003ewhere \u003cem\u003eY\u003c/em\u003e\u003csub\u003ei\u003c/sub\u003e represents travel intention for video \u003cem\u003ei\u003c/em\u003e, \u003cem\u003eF\u003c/em\u003e\u003csub\u003ei1\u003c/sub\u003e and \u003cem\u003eF\u003c/em\u003e\u003csub\u003ei2\u003c/sub\u003e denote fused feature statistics (mean activation and temporal consistency respectively), \u003cem\u003eC\u003c/em\u003e\u003csub\u003ei1\u003c/sub\u003e and \u003cem\u003eC\u003c/em\u003e\u003csub\u003ei2\u003c/sub\u003e represent cognitive load measures (intrinsic content complexity and extraneous modal interference respectively), \u003cem\u003eD\u003c/em\u003e\u003csub\u003ei1\u003c/sub\u003e and \u003cem\u003eD\u003c/em\u003e\u003csub\u003ei2\u003c/sub\u003e are dummy variables for video categories (with Destination Attraction as reference category), and ε\u003csub\u003ei\u003c/sub\u003e is the error term. Given identified assumption violations, robust standard errors were computed to ensure valid inference. Temporal causality among modalities was examined using Granger causality tests with a maximum lag of five windows, preceded by Augmented Dickey-Fuller tests to verify stationarity.\u003c/p\u003e\u003cp\u003eRecognizing that digital influence mechanisms may operate through non-linear pathways, we implemented Random Forest regression to capture complex relationships and interactions. To prevent overfitting and ensure generalizability, feature selection was performed exclusively on the training set (75% of data), removing features with correlations exceeding 0.98 to eliminate redundancy. Model hyperparameters were conservatively configured (100 estimators, max_depth\u0026thinsp;=\u0026thinsp;8, min_samples_split\u0026thinsp;=\u0026thinsp;10, max_features='sqrt') to balance complexity with generalization. Performance evaluation employed 5-fold cross-validation on the training set and held-out test set validation, with bootstrap resampling (1,000 iterations) providing confidence intervals for performance metrics. Model interpretation utilized SHAP (SHapley Additive exPlanations) analysis to decompose predictions into feature contributions, revealing both global importance patterns and local decision boundaries.\u003c/p\u003e\u003cp\u003eAll analyses were implemented in Python 3.9 using scikit-learn (1.3.0) for machine learning, statsmodels (0.14.0) for statistical testing, PyTorch (2.0.1) for deep learning components, and SHAP (0.42.1) for model interpretation. Computations were performed on hardware featuring NVIDIA RTX 3090 GPU and 64GB RAM to ensure computational efficiency.\u003c/p\u003e\u003c/div\u003e"},{"header":"4. Results","content":"\u003cdiv id=\"Sec15\" class=\"Section2\"\u003e\u003ch2\u003e4.1 Data-Driven Categorization of Tourism Video Content\u003c/h2\u003e\u003cp\u003eThe computational analysis of 2,650 tourism-themed videos from Bilibili through BERTopic modeling revealed a hierarchical structure of 67 distinct topics, subsequently consolidated into three primary categories via agglomerative clustering with Ward's linkage method. The hierarchical clustering dendrogram (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e) demonstrated clear semantic boundaries between topic clusters, enabling robust categorization into: Destination Attraction videos (37.4%, n\u0026thinsp;=\u0026thinsp;991), Tourism Narrative Communication content (42.1%, n\u0026thinsp;=\u0026thinsp;1,116), and Overseas Cultural Experience materials (20.5%, n\u0026thinsp;=\u0026thinsp;543) (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eDistribution and Characteristics of Video Categories\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"5\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eCategory\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eN\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003e%\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eRepresentative Topics\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eKey Lexical Features\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDestination Attraction\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e991\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e37.4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\"Qingdao Seaside Travel\", \"Hokkaido Winter Snow\", \"Tibet Self-Driving\"\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eGeographic names, visual descriptors, spatial terms\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eTourism Narrative Communication\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e1,116\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e42.1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\"Leisure Travel Log\", \"Youth Group Tour\", \"Travel Guide Alone\"\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003ePersonal pronouns, temporal markers, emotional language\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eOverseas Cultural Experience\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e543\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e20.5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\"Travel News in India\", \"Southern Europe Vlog\", \"Korean Girls Travel to Shanghai\"\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eCultural terms, international place names, comparative language\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eThe validity of this categorization was quantitatively confirmed through similarity matrix analysis (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e), which demonstrated a pronounced block-diagonal structure. Within-category similarity scores (M\u0026thinsp;=\u0026thinsp;0.82, SD\u0026thinsp;=\u0026thinsp;0.07) significantly exceeded between-category similarities (M\u0026thinsp;=\u0026thinsp;0.64, SD\u0026thinsp;=\u0026thinsp;0.11), yielding a substantial effect size (t(2648)\u0026thinsp;=\u0026thinsp;47.28, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001, Cohen's d\u0026thinsp;=\u0026thinsp;1.84), confirming genuine categorical distinctions rather than arbitrary divisions.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eThe intertopic distance map (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e) further corroborated categorical distinctions, with each category occupying distinct regions in the reduced dimensional space. Temporal analysis from 2018 to 2024 (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e) revealed differential growth trajectories: Overseas Cultural Experience demonstrated the most pronounced expansion (CAGR\u0026thinsp;=\u0026thinsp;35.2%), particularly accelerating post-2023 coinciding with relaxed international travel restrictions, while Destination Attraction showed steady growth (CAGR\u0026thinsp;=\u0026thinsp;11.2%).\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec16\" class=\"Section2\"\u003e\u003ch2\u003e4.2 Multimodal and Cognitive Profiles of Video Categories\u003c/h2\u003e\u003cp\u003eTo test H1, we conducted systematic analyses of variance across 102 computationally extracted features. The analysis revealed that 94 of 102 features (92.2%) demonstrated significant categorical variation at p\u0026thinsp;\u0026lt;\u0026thinsp;0.05 after Benjamini-Hochberg correction for multiple comparisons (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eSummary of ANOVA Results Across Feature Domains\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"5\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eFeature Domain\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eN Features\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eSignificant (%)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eMean η\u0026sup2;\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eMax η\u0026sup2;\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAudio Features\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e89\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e92.1%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.044\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.072\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eVisual Features\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e8\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e87.5%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.031\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.052\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDanmaku Interaction\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e100%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.048\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.063\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eCognitive Load*\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e-\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e-\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e-\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003ctfoot\u003e\u003ctr\u003e\u003ctd colspan=\"5\"\u003eNote: While Levene's tests indicated heterogeneous variances for some features (p\u0026thinsp;\u0026lt;\u0026thinsp;0.05), the large effect sizes and consistent patterns across multiple features support the robustness of our categorical distinctions. Welch's ANOVA confirmed all significant findings.\u003c/td\u003e\u003c/tr\u003e\u003c/tfoot\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eDanmaku interaction patterns revealed striking categorical differences (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e). Destination Attraction videos elicited the highest commentary density (M\u0026thinsp;=\u0026thinsp;1.50 per window, SD\u0026thinsp;=\u0026thinsp;5.89), significantly exceeding Overseas Cultural Experience videos (M\u0026thinsp;=\u0026thinsp;0.73, SD\u0026thinsp;=\u0026thinsp;4.85; F(2, 2647)\u0026thinsp;=\u0026thinsp;89.34, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001, η\u0026sup2; = 0.063). This pattern suggests differentiated audience engagement strategies, with scenic content triggering immediate reactions while cross-cultural content promotes contemplative viewing.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eComprehensive acoustic profiling across 89 features revealed distinctive sonic fingerprints for each category (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e). Spectral centroid analysis demonstrated that Destination Attraction videos contained the brightest audio (M\u0026thinsp;=\u0026thinsp;1002.19 Hz), while MFCC analysis showed Overseas Cultural Experience videos had substantially lower spectral energy (MFCC 0: M = -175.48), suggesting distinct production contexts.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eVisual analysis revealed that Tourism Narrative Communication videos contained significantly more detected objects (M\u0026thinsp;=\u0026thinsp;32.40, SD\u0026thinsp;=\u0026thinsp;49.11) compared to other categories (F(2, 2647)\u0026thinsp;=\u0026thinsp;52.78, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001, η\u0026sup2; = 0.038), while CLIP embedding magnitudes indicated Overseas Cultural Experience videos possessed richer semantic content (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003e).\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec17\" class=\"Section2\"\u003e\u003ch2\u003e4.3 Limitations of Linear Modeling: An Exploratory Analysis\u003c/h2\u003e\u003cp\u003e\u003cstrong\u003eImportant Note on Statistical Assumptions\u003c/strong\u003e\u003cp\u003eBefore presenting the linear regression results, we must acknowledge severe violations of statistical assumptions that fundamentally limit their interpretation. Diagnostic tests revealed\u003c/p\u003e\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eExtreme non-normality (Jarque-Bera\u0026thinsp;=\u0026thinsp;7,742,617, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001; skewness\u0026thinsp;=\u0026thinsp;11.89, kurtosis\u0026thinsp;=\u0026thinsp;266.74)\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003e Severe multicollinearity with VIF values exceeding 100 for some features (condition number\u0026thinsp;=\u0026thinsp;487)\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eInfluential observations (maximum Cook's distance\u0026thinsp;=\u0026thinsp;0.2179)\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eThese violations render traditional linear inference unreliable. We present these results solely for completeness and as motivation for the non-linear approaches that follow. The linear model's minimal explanatory power (R\u0026sup2; = 0.033, Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e) further confirms the inadequacy of linear frameworks for these data. Regression diagnostic plots (Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e) visually confirm these severe assumption violations..\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eMultiple Regression Results for Travel Intention\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"6\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\"\u0026minus;\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eVariable\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eβ\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eSE\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003et\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003ep\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003e95% CI\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eConstant\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e9.79\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e44.09\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.22\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.824\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\"\u0026minus;\" colname=\"c6\"\u003e\u003cp\u003e[-76.66, 96.24]\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eFused Temporal Consistency\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e-5.20\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.80\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e-6.50\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e\u0026lt;\u0026thinsp;0.001***\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\"\u0026minus;\" colname=\"c6\"\u003e\u003cp\u003e[-6.77, -3.63]\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eCategory: Overseas Cultural Experience\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e-4.59\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.91\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e-5.06\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e\u0026lt;\u0026thinsp;0.001***\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\"\u0026minus;\" colname=\"c6\"\u003e\u003cp\u003e[-6.37, -2.81]\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eFused Statistical Mean Activation\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e-46.70\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e31.85\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e-1.47\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.143\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\"\u0026minus;\" colname=\"c6\"\u003e\u003cp\u003e[-109.15, 15.75]\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eCognitive Intrinsic Load Content Complexity\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e13.43\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e37.55\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.36\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.721\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\"\u0026minus;\" colname=\"c6\"\u003e\u003cp\u003e[-60.20, 87.06]\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eCognitive Extraneous Load Modal Interference\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e3.07\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e9.04\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.34\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.734\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\"\u0026minus;\" colname=\"c6\"\u003e\u003cp\u003e[-14.65, 20.79]\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eCategory: Tourism Narrative Communication\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e1.13\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e1.26\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.90\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.367\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\"\u0026minus;\" colname=\"c6\"\u003e\u003cp\u003e[-1.33, 3.60]\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eNote: ***p\u0026thinsp;\u0026lt;\u0026thinsp;0.001; Results should be interpreted as exploratory patterns only due to assumption violations.\u003c/p\u003e\u003cp\u003eFigure \u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e: Regression Diagnostic Plots\u003c/p\u003e\u003cp\u003eDespite these limitations, Granger causality tests revealed potentially interesting temporal patterns. For Overseas Cultural Experience videos (5,430 windows analyzed), visual changes preceded audio responses in 14.3% of windows (p\u0026thinsp;\u0026lt;\u0026thinsp;0.05), which subsequently triggered Danmaku commentary in 13.5% of windows, creating a visual-to-audio-to-text cascade (χ\u0026sup2; = 287.43, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001).\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec18\" class=\"Section2\"\u003e\u003ch2\u003e4.4 Non-linear Modeling and Feature Importance Analysis\u003c/h2\u003e\u003cp\u003eGiven the failure of linear approaches, we implemented Random Forest regression with careful attention to overfitting concerns. Systematic ablation studies across feature configurations revealed a striking pattern: cognitive features dramatically outperformed traditional multimodal features in predicting travel intention (Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e).\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eAblation Study Results\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"7\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eModel Configuration\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eN Features\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eTest R\u0026sup2;\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eTrain R\u0026sup2;\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eOverfitting Gap\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eMAE\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003eInterpretation\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eCognitive Features Only\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.518\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.948\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.430*\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e2.10\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eHigh risk of overfitting\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAll Multimodal Features\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e48\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.366\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.822\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.456*\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e4.74\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eSevere overfitting\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eText Features Only\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e7\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.144\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.525\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.381*\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e7.71\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eModerate overfitting\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eFused Features Only\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.084\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.436\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.352*\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e8.47\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eModerate overfitting\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eVisual Features Only\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.078\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.418\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.340*\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e8.52\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eModerate overfitting\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAudio Features Only\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e10\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.039\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.382\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.343*\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e9.47\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eModerate overfitting\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e*Note: Overfitting gaps\u0026thinsp;\u0026gt;\u0026thinsp;0.3 indicate substantial overfitting, requiring cautious interpretation.\u003c/p\u003e\u003cp\u003eThe model complexity versus performance analysis (Fig.\u0026nbsp;\u003cspan refid=\"Fig9\" class=\"InternalRef\"\u003e9\u003c/span\u003e) revealed while the cognitive features model achieved the highest test R\u0026sup2; (0.518), the substantial train-test gap (0.430) suggests the model may be capturing idiosyncratic patterns rather than generalizable relationships. This finding supports H3a (cognitive features dominate prediction) but with important caveats regarding generalizability.\u003c/p\u003e\u003cp\u003eInterestingly, the low canonical correlation between audio and visual features (CCA\u0026thinsp;=\u0026thinsp;0.028) suggests weak inherent alignment in PUGC content, further supporting our finding that cognitive processing rather than multimodal harmony drives travel intention formation. This weak audiovisual correlation contradicts assumptions about professional content quality and indicates that PUGC creators may not optimize cross-modal coherence, making cognitive interpretation by viewers the critical determinant of influence.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eSHAP analysis revealed that 'cognitive intrinsic load conceptual difficulty' dominated with 40.1% of total importance, followed by 'text temporal variability' (11.9%) and 'fused temporal consistency' (8.4%). The concentration of importance in cognitive features (42.2% of importance from 10.4% of features) suggests that travel intention formation depends primarily on cognitive processing rather than sensory richness, though this interpretation must be tempered by overfitting concerns.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eSHAP dependence analysis (Fig.\u0026nbsp;\u003cspan refid=\"Fig11\" class=\"InternalRef\"\u003e11\u003c/span\u003e) revealed complex non-linear relationships with clear threshold effects and saturation patterns, explaining why linear models failed. These non-monotonic patterns, while compelling, require validation on independent datasets given the overfitting indicators.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eThe final optimized model's performance (Fig.\u0026nbsp;\u003cspan refid=\"Fig12\" class=\"InternalRef\"\u003e12\u003c/span\u003e) demonstrates the final optimized model, using conservative hyperparameters to reduce overfitting, achieved Test R\u0026sup2; = 0.378 (95% Bootstrap CI: [0.249, 0.507]) with MAE\u0026thinsp;=\u0026thinsp;4.33. While showing lower R\u0026sup2; than the pure cognitive model, this configuration demonstrated a more acceptable train-test gap (0.31), suggesting better generalizability.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e"},{"header":"5. Discussion","content":"\u003cp\u003eThis investigation reveals a fundamental reconceptualization of digital tourism influence mechanisms within platform-mediated environments. Analysis of 2,650 PUGC videos uncovers three transformative insights that challenge conventional understanding of tourism marketing in digital contexts, as summarized in Table\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e. Cognitive processing demonstrates primacy over sensory stimulation in determining travel intention formation, with cognitive features dominating predictive importance despite comprising a small fraction of analyzed features. Furthermore, multimodal influence operates through non-linear threshold mechanisms rather than additive effects, evidenced by the dramatic performance disparity between linear and non-linear models. Additionally, synchronized social commentary through Danmaku functions as cognitive scaffolding, transforming passive viewing into collective sense-making processes with distinct categorical patterns.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eSummary of Hypothesis Testing Results\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"4\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eHypothesis\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eResult\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eKey Evidence\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eTheoretical Implications\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eH1a\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eStrongly Supported\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e92.2% of features showed significant variation (p\u0026thinsp;\u0026lt;\u0026thinsp;0.05)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003ePUGC categories represent distinct multimodal communication strategies\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eH1b\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eSupported\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eCognitive features showed significant categorical differences\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eContent complexity systematically varies across tourism communication goals\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eH2a\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNot Supported\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eLinear model R\u0026sup2; = 0.033; assumption violations\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eLinear frameworks fundamentally inadequate for PUGC data\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eH2b\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eSupported\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eGranger causality confirmed (χ\u0026sup2; = 287.43, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003ePredictable influence cascades exist across modalities\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eH3a\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003ePartially Supported\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eNon-linear R\u0026sup2; = 0.378 (optimized) vs. linear R\u0026sup2; = 0.033; overfitting concerns with higher R\u0026sup2; models\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eNon-linear integration improves prediction but generalizability remains challenging\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eH3b\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eStrongly Supported\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eCognitive features 42.2% importance with 10.4% of features\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eCognitive processing and interaction dominate over content features\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003ctfoot\u003e\u003ctr\u003e\u003ctd colspan=\"4\"\u003eNote: H3a's R\u0026sup2; = 0.518 represents the cognitive-only model with substantial overfitting (train-test gap\u0026thinsp;=\u0026thinsp;0.430). The final optimized model achieved R\u0026sup2; = 0.378 with reduced overfitting.\u003c/td\u003e\u003c/tr\u003e\u003c/tfoot\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cdiv id=\"Sec20\" class=\"Section2\"\u003e\u003ch2\u003e5.1 The Cognitive Primacy Revolution in Digital Tourism Influence Mechanisms\u003c/h2\u003e\u003cp\u003eThe sixfold predictive advantage of cognitive features (R\u0026sup2; = 0.518) relative to audiovisual features (R\u0026sup2; \u0026lt; 0.08) challenges foundational assumptions underlying tourism marketing since Urry's (1990)(Urry 1990) conceptualization of the tourist gaze, specifically the presumption that visual stimulation drives travel desire. Although observed overfitting (train-test gap\u0026thinsp;=\u0026thinsp;0.430) necessitates interpretive caution, the pattern's consistency across multiple model specifications suggests genuine phenomena whereby cognitive engagement mediates and potentially supersedes sensory appeal within interactive digital contexts.\u003c/p\u003e\u003cp\u003eThis finding acquires enhanced theoretical significance when considered alongside methodological innovations employed in this investigation. Traditional survey methodologies lack capacity to detect micro-temporal cognitive dynamics characterizing digital content consumption(Otto et al. 2022). Through computational operationalization of cognitive load via multimodal feature complexity, cross-modal interference patterns, and semantic density variations, previously invisible influence mechanisms emerge. The concentration of predictive power, wherein 42.2% of model importance derives from merely 10.4% of features, suggests mental activation represents not simply another factor but potentially the master mechanism determining whether content translates into behavioral intention.\u003c/p\u003e\u003cp\u003eThe remarkably weak canonical correlation between audio and visual features (CCA\u0026thinsp;=\u0026thinsp;0.028) provides crucial corroborating evidence. In contrast to professional content where audiovisual elements achieve careful synchronization, PUGC creators appear to prioritize authenticity over production coherence. This production heterogeneity paradoxically increases cognitive demands, compelling viewers toward active sense-making rather than passive consumption. The implications prove profound, suggesting that amateur aesthetics characteristic of PUGC achieve effectiveness not despite but because of cognitive demands imposed, thereby transforming viewers from spectators into mental participants.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec21\" class=\"Section2\"\u003e\u003ch2\u003e5.2 The Overseas Cultural Experience Paradox and Virtual Travel Substitution\u003c/h2\u003e\u003cp\u003eAmong our findings, the negative association between Overseas Cultural Experience videos and travel intention (β = -4.59, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001) presents the most theoretically provocative result. This counterintuitive pattern, wherein rich cultural content diminishes rather than stimulates travel desire, demands careful theoretical consideration through multiple explanatory frameworks.\u003c/p\u003e\u003cp\u003eThe mediated substitution hypothesis advanced by Liu et al. (2023)(Liu et al. 2023) offers one explanatory pathway, proposing that immersive virtual experiences satisfy psychological needs motivating travel, including novelty seeking, cultural learning, and social distinction, without requiring physical displacement. Our data provides supporting evidence through this category's distinctive profile, characterized by highest semantic richness evidenced in CLIP embedding magnitude, lowest Danmaku density (M\u0026thinsp;=\u0026thinsp;0.73), and moderate cognitive load levels. This combination suggests contemplative consumption patterns wherein viewers achieve cultural gratification through viewing alone.\u003c/p\u003e\u003cp\u003eNevertheless, alternative mechanisms warrant careful consideration. Selection bias potentially explains these patterns if audiences consuming extensive cultural content comprise primarily armchair travelers seeking cultural knowledge without corresponding travel intention. Alternatively, cognitive overload mechanisms, supported by non-linear models revealing threshold effects(Hossain and Yeasin 2017), suggest excessive cultural complexity might overwhelm rather than inspire, particularly when viewers lack appropriate cultural scaffolding for interpretation. Moreover, social proof deficit, evidenced by low Danmaku density, removes collective enthusiasm that typically transforms individual interest into shared aspiration.\u003c/p\u003e\u003cp\u003eMost compelling, temporal analysis reveals Overseas Cultural Experience videos demonstrate strongest Granger causality patterns, with visual elements preceding audio responses that subsequently generate textual commentary in 14.3% of analytical windows. This sophisticated orchestration paradoxically may undermine travel intention by providing such comprehensive virtual experiences that physical travel appears redundant.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec22\" class=\"Section2\"\u003e\u003ch2\u003e5.3 Non-linear Dynamics and Optimal Cognitive Load Configurations\u003c/h2\u003e\u003cp\u003eThe stark performance disparity between linear models (R\u0026sup2; = 0.033) and non-linear approaches (R\u0026sup2; = 0.378, optimized) provides empirical validation for threshold-based influence mechanisms. SHAP analysis reveals specific non-linearities wherein cognitive load demonstrates inverted-U relationships with travel intention, achieving peak influence at moderate complexity levels (standardized value\u0026thinsp;\u0026asymp;\u0026thinsp;0.4) before subsequent decline. While this pattern aligns with flow theory conceptualized by Csikszentmihalyi (1990)(Mihaly Csikszentmihalyi 1990), manifestation differs substantially between tourism contexts and traditional learning environments.\u003c/p\u003e\u003cp\u003eCategory-specific optimal zones discovered through systematic analysis provide actionable insights for content optimization. Destination Attraction content maximizes influence at lower cognitive load levels (0.2\u0026ndash;0.3 standardized units) combined with elevated social density (1.50 comments per window), suggesting viewers seek social validation for aesthetic experiences. Conversely, Tourism Narrative Communication achieves optimal performance at higher cognitive load levels (0.5\u0026ndash;0.6 standardized units), indicating audiences expect and reward storytelling complexity. Meanwhile, Overseas Cultural Experience demonstrates the flattest response curve, suggesting cognitive load may prove less determinative than cultural authenticity markers for this content category.\u003c/p\u003e\u003cp\u003eThese patterns necessitate fundamental reconceptualization of digital tourism influence. Rather than maximizing production quality or minimizing cognitive effort, effective PUGC operates within category-specific cognitive configurations where challenge activates without overwhelming processing capacity.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec23\" class=\"Section2\"\u003e\u003ch2\u003e5.4 Methodological Innovation and Theoretical Advancement Integration\u003c/h2\u003e\u003cp\u003eOur computational approach transcends mere efficiency gains in measuring established phenomena, instead revealing influence mechanisms invisible to traditional methodologies. Extraction of 102 objective features capturing millisecond-level multimodal dynamics enabled discovery of three critical phenomena previously unobservable.\u003c/p\u003e\u003cp\u003eFirst, attention cascades demonstrate predictable operation across modalities, wherein visual changes trigger audio responses that subsequently generate Danmaku commentary. This temporal choreography, detected through Granger causality analysis, suggests successful creators intuitively orchestrate influence sequences that traditional cross-sectional analysis(Spector 2019) cannot detect.\u003c/p\u003e\u003cp\u003eSecond, cognitive interference patterns between modalities reveal conditions under which information streams compete versus complement. The negative temporal consistency coefficient (β = -5.20) indicates monotony diminishes engagement, whereas controlled variation maintains cognitive activation, a finding achievable only through computational temporal analysis.\u003c/p\u003e\u003cp\u003eThird, collective sense-making dynamics emerge through Danmaku density bursts at specific narrative moments, revealing transformation points where individual viewing becomes social experience. These crystallization points, characterized by concentrated travel intention expressions, provide blueprints for engineering viral influence moments.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec24\" class=\"Section2\"\u003e\u003ch2\u003e5.5 Tourism Marketing Transformation from Broadcasting to Cognitive Choreography\u003c/h2\u003e\u003cp\u003ePractical implications emerging from our findings necessitate fundamental reconsideration of tourism marketing strategies. Traditional paradigms emphasizing inspiration through beauty, information through features, and persuasion through benefits assume passive viewers processing additive information. Conversely, our evidence indicates active viewers navigate complex information ecosystems wherein cognitive engagement determines influence outcomes.\u003c/p\u003e\u003cp\u003eDestination marketing organizations must therefore shift emphasis from production quality toward cognitive design. Rather than maximizing visual appeal, content should optimize cognitive load within category-specific configurations(Bai et al. 2025). Destination attraction content should maintain visual stability while maximizing social interaction opportunities(Malafouris et al. 2021). Tourism narratives require strategic multimodal variation sustaining engagement across extended viewing periods(Mattei 2023). Cultural content must carefully balance immersion with aspiration, avoiding substitution effects wherein virtual satisfaction replaces travel motivation(Hu et al. 2023).\u003c/p\u003e\u003cp\u003eContent creators receive validation for PUGC approaches while gaining optimization guidance. Weak audiovisual correlation (CCA\u0026thinsp;=\u0026thinsp;0.028) suggests production perfection may prove counterproductive, as authentic heterogeneity engages more effectively than polished coherence. Creators should therefore focus on cognitive choreography through strategic information revelation, controlled complexity escalation, and social interaction facilitation(Sigala 2016).\u003c/p\u003e\u003cp\u003ePlatform designers encounter algorithmic innovation opportunities suggested by our results. Current recommendation systems optimizing view duration or engagement metrics may overlook cognitive activation patterns driving behavioral intention(Wen et al. 2019). Platforms could develop creator tools facilitating attention cascade orchestration, optimal Danmaku density maintenance, and cognitive load monitoring capabilities(He and Tang 2017; Yang 2020b; Peng et al. 2025).\u003c/p\u003e\u003c/div\u003e"},{"header":"6. Limitations and Future Research Directions","content":"\u003cp\u003eWhile this investigation advances understanding of digital tourism influence mechanisms, several methodological and theoretical boundaries illuminate both constraints and opportunities for future inquiry.\u003c/p\u003e\u003cp\u003eOur reliance on comment-based intention measures captures expressed enthusiasm among 8.7% of viewers rather than actual booking behaviors, representing the most fundamental limitation. Paradoxically, this constraint strengthens theoretical contributions, as cognitive mechanisms dominating even expressed intention likely play stronger roles in actual behavior requiring greater commitment. Future research should establish conversion pathways through platform-DMO partnerships enabling behavioral validation. Similarly, single-platform data from Bilibili constrains cross-cultural generalizability while providing exceptional ecosystem depth. Bilibili's distinctive Danmaku feature, absent from Western platforms, enabled unique insights into synchronized social viewing. Comparative analysis across YouTube, TikTok, and Instagram Reels would distinguish universal cognitive mechanisms from platform-specific affordances, particularly revealing whether platforms lacking Danmaku exhibit stronger audiovisual effects due to reduced attentional competition.\u003c/p\u003e\u003cp\u003eThe substantial overfitting observed in certain models, with train-test gaps reaching 0.430, raises critical questions about cognitive feature stability. Rather than purely methodological weakness, this instability may carry theoretical significance, suggesting cognitive influence mechanisms are inherently context-dependent and require adaptive rather than universal modeling approaches. Whether overfitting reflects genuine cognitive complexity or measurement limitations remains an empirical question demanding resolution.\u003c/p\u003e\u003cp\u003eThese limitations nonetheless establish several theoretical frontiers warranting systematic investigation. The cognitive efficiency paradox, wherein five cognitive features outperform 22 audiovisual features, suggests information compression mechanisms extract essential decision inputs through either evolutionary spatial adaptations or learned digital heuristics. Equally intriguing, the authenticity-chaos hypothesis emerges from remarkably weak audiovisual correlation (CCA\u0026thinsp;=\u0026thinsp;0.028), contradicting media richness theory while suggesting production incoherence paradoxically signals authenticity. Experimental manipulation of production coherence could establish causal relationships and boundary conditions for this counterintuitive phenomenon.\u003c/p\u003e\u003cp\u003eMost provocatively, the substitution threshold model implied by Overseas Cultural Experience findings indicates certain content satisfies rather than stimulates travel desire. As virtual reality tourism emerges, identifying precise thresholds between inspiration and substitution becomes critical for industry sustainability. Research mapping psychological needs satisfied through virtual versus physical travel could inform strategic responses to technological disruption. Meanwhile, the collective intelligence mechanism manifested through Danmaku raises fundamental questions about social information processing. Eye-tracking studies could determine whether viewers process synchronized commentary as informational input or atmospheric enhancement, distinguishing wisdom of crowds from conformity pressure.\u003c/p\u003e\u003cp\u003eOur findings arrive at a critical juncture as tourism marketing confronts AI-generated content, virtual reality experiences, and intensifying attention economy dynamics. Three implications emerge with particular force. First, cognitive engagement may constitute the emerging scarce resource in tourism marketing, as infinite AI-generated content confronts finite human processing capacity. Second, platform mediation will intensify given evidence that affordances fundamentally alter influence mechanisms, necessitating platform-specific rather than universal strategies. Third, boundaries between virtual and physical tourism will progressively blur, with our Overseas Cultural Experience findings previewing direct competition between mediated and embodied experiences. Tourism's value proposition must therefore evolve beyond visual consumption toward experiences that resist virtualization.\u003c/p\u003e\u003cp\u003eThese considerations suggest tourism marketing stands at an inflection point where traditional paradigms yield to cognitive-interactive frameworks. Success will increasingly depend on understanding and orchestrating mental engagement rather than maximizing sensory stimulation, a transformation our findings both document and accelerate.\u003c/p\u003e"},{"header":"7. Conclusion","content":"\u003cp\u003eThis investigation illuminates a fundamental reorientation in digital tourism influence, wherein cognitive activation supersedes sensory persuasion as the primary mechanism driving travel intention formation. Computational analysis of 2,650 PUGC videos from Bilibili reveals that cognitive engagement accounts for substantially greater variance in behavioral intention than audiovisual attributes, challenging orthodox assumptions underlying tourism marketing practice.\u003c/p\u003e\u003cp\u003eThree theoretical contributions emerge from this analysis. First, the identification of non-linear, threshold-based influence patterns explains the persistent failure of linear frameworks in digital contexts, demonstrating that cognitive features generate predictive power through discontinuous rather than additive relationships. Second, Danmaku's transformation of solitary viewing into collective sense-making constitutes a novel social influence pathway, wherein synchronized commentary creates cognitive scaffolding absent from traditional media consumption. Third, the counterintuitive finding that culturally rich content diminishes rather than enhances travel intention illuminates an emerging substitution dynamic, presaging fundamental challenges as virtual experiences increasingly approximate physical travel's psychological rewards.\u003c/p\u003e\u003cp\u003eMethodologically, this study validates computational approaches for detecting influence mechanisms beyond the reach of conventional methods. Extraction and analysis of 102 objective features capturing millisecond-level multimodal dynamics revealed attention cascades, cognitive interference patterns, and collective crystallization moments invisible to survey-based inquiry. Such granular temporal resolution proves essential for understanding how influence unfolds through orchestrated sequences rather than static attributes.\u003c/p\u003e\u003cp\u003eThese findings acquire heightened significance as tourism marketing navigates technological disruption through AI-generated content and immersive virtual experiences. The primacy of cognitive engagement over production quality suggests competitive advantage will increasingly derive from mental activation rather than sensory sophistication. Content creators who master cognitive choreography\u0026mdash;the strategic orchestration of complexity, variation, and social interaction\u0026mdash;will shape travel decisions more effectively than those pursuing traditional aesthetic excellence. Thus, the pathway from watching to wishing, this investigation reveals, traverses not the eye but the mind, positioning cognitive design as tourism marketing's emerging imperative.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eF.Y. designed the study, developed the methodology, implemented software, conducted the investigation, and prepared the original manuscript draft.M.Y. performed data curation, validation, and contributed to reviewing and editing the manuscript.S.S. provided supervision and contributed to the critical review of the manuscript.X.W. provided supervision and project oversight and contributed to manuscript revision.All authors reviewed and approved the final manuscript.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eAlaei AR, Becken S, Stantic B (2019) Sentiment Analysis in Tourism: Capitalizing on Big Data. Journal of Travel Research 58:175\u0026ndash;191. https://doi.org/10.1177/0047287517747753\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAnh HH (2024) Harmonizing Engagement: The Impact of Music on Consumer Interaction in Social Media Marketing. IJEBMR 08:182\u0026ndash;202. https://doi.org/10.51505/IJEBMR.2024.81013\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAsyraff MA, Hanafiah M, Aminuddin M, Mahdzar N (2023) Adoption of the stimulus-organism-response (S-O-R) model in hospitality and tourism research: systematic literature review and future research directions adoption of the stimulus-organism-response (S-O-R) model in hospitality and tourism research: systematic literature review and future research directions. Asia-pac J Innov Hosp Tour (APJIHT) 12:2023\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eB. Joseph Pine II, Gilmore James H. (1998) Welcome to the Experience Economy. Harvard Business Review\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBai S, Li Z, He H, Fan W (2025) Engaging tourists from city aesthetics: evidence from multi-modal analysis of computer vision and text mining. APJML. https://doi.org/10.1108/APJML-02-2025-0287\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBailenson J, Guadagno R, Aharoni E, et al (2004) Comparing behavioral and self-report measures of embodied agents\u0026rsquo; social presence in immersive virtual environments\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBaltrusaitis T, Ahuja C, Morency L-P (2019) Multimodal Machine Learning: A Survey and Taxonomy. IEEE Trans Pattern Anal Mach Intell 41:423\u0026ndash;443. https://doi.org/10.1109/TPAMI.2018.2798607\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBayerl PS, Paul KI (2011) What Determines Inter-Coder Agreement in Manual Annotations? A Meta-Analytic Investigation. Computational Linguistics 37:699\u0026ndash;725. https://doi.org/10.1162/COLI_a_00074\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003e(Brian) Lin M-T, Zhu D, Liu C, Kim PB (2022) A systematic review of empirical studies of pro-environmental behavior in hospitality and tourism contexts. IJCHM 34:3982\u0026ndash;4006. https://doi.org/10.1108/IJCHM-12-2021-1478\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBrunken R, Plass JL, Leutner D (2003) Direct Measurement of Cognitive Load in Multimedia Learning. Educational Psychologist 38:53\u0026ndash;61. https://doi.org/10.1207/S15326985EP3801_7\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBuckingham D (2009) A Commonplace Art? Understanding Amateur Media Production. Video Cultures 23\u0026ndash;50. https://doi.org/10.1057/9780230244696_2\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCauchard JR, Gover W, Chen WL, et al (2024) Orality, multimodality and creativity in digital writing: Chinese users\u0026rsquo; experiences and practices with bullet comments on \u003cem\u003eBilibili\u003c/em\u003e. Social Semiotics 34:368\u0026ndash;394. https://doi.org/10.1080/10350330.2022.2120387\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCaulley DN (1994) Review \u0026amp; Booknote: The Unobtrusive Researcher: A Guide to Methods. Media Information Australia 73:118\u0026ndash;118. https://doi.org/10.1177/1329878X9407300133\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChemin M, Silva CP, Vikou SVP (2025) User-generated content (UGC) in tourist attractions and destinations: systematic literature review and perspectives for management\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChen M (1986) The Impact of Authenticity and Credibility Factors on Consumer Behavior in the Sustainable Fashion Industry on Weibo\u0026ndash;Taking Micro-Influencers and Mega-Influencers as Examples. Tourism Management 91:322\u0026ndash;328. https://doi.org/10.54254/2754-1169/91/20241077\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChen Y, Gao Q, Rau P-LP (2015) Understanding Gratifications of Watching Danmaku Videos \u0026ndash; Videos with Overlaid Comments. In: Rau PLP (ed) Lecture Notes in Computer Science. Springer International Publishing, Cham, pp 153\u0026ndash;163\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChen Z (2025) Theoretical development of the tourist experience: a future perspective. Tourism Recreation Research 50:199\u0026ndash;213. https://doi.org/10.1080/02508281.2023.2255939\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCialdini R (1993) Influence: Science and Practice\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDiehl M, Wahl H-W, Freund A (2017) Ecological Validity as a Key Feature of External Validity in Research on Human Development. Research in Human Development 14:177\u0026ndash;181. https://doi.org/10.1080/15427609.2017.1340053\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDimitriou CK, AbouElgheit E (2019) Understanding generation Z\u0026rsquo;s travel social decision-making. Tour hosp manag 25:311\u0026ndash;334. https://doi.org/10.20867/thm.25.2.4\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eEhlers J (2020) Exploring the effect of transient cognitive load on bodily arousal and secondary task performance. In: Proceedings of Mensch und Computer 2020. ACM, New York, NY, USA, pp 7\u0026ndash;10\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFan M, Ostic D, Han K, Shar S (2023) The Impact of Danmaku Information Quality on Consumers\u0026rsquo; Impulsive Consumption Behavior. Proceedings 2023:10474. https://doi.org/10.5465/AMPROC.2023.10474abstract\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFang J, Chen L, Wen C, Prybutok VR (2018) Co-viewing Experience in Video Websites: The Effect of Social Presence on E-Loyalty. International Journal of Electronic Commerce 22:446\u0026ndash;476. https://doi.org/10.1080/10864415.2018.1462929\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFazli-Salehi R, Jahangard M, Torres IM, et al (2022) Social media reviewing channels: the role of channel interactivity and vloggers\u0026rsquo; self-disclosure in consumers\u0026rsquo; parasocial interaction. JCM 39:242\u0026ndash;253. https://doi.org/10.1108/JCM-06-2020-3866\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFilieri R, Christodoulides G, Nicolau JL (2025) Emerging Sources, Formats, Channels, Devices, and Audiences in Modern eWOM Communications. Psychology and Marketing 42:2430\u0026ndash;2443. https://doi.org/10.1002/mar.22239\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFotis J, Buhalis D, Rossides N (2012) Social Media Use and Impact during the Holiday Travel Planning Process. Information and Communication Technologies in Tourism 2012 13\u0026ndash;24. https://doi.org/10.1007/978-3-7091-1142-0_2\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGong W, Li X (2017) Engaging fans on microblog: the synthetic influence of parasocial interaction and source characteristics on celebrity endorsement\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGretzel U (2018) Influencer Marketing in Travel and Tourism\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGrimmer J, Stewart BM (2013) Text as Data: The Promise and Pitfalls of Automatic Content Analysis Methods for Political Texts. Polit anal 21:267\u0026ndash;297. https://doi.org/10.1093/pan/mps028\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGrootendorst M (2020) BERTopic: Neural topic modeling with a class-based TF-IDF procedure\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHaji FA, Rojas D, Childs R, et al (2015) Measuring cognitive load: performance, mental effort and simulation task complexity. Medical Education 49:815\u0026ndash;827. https://doi.org/10.1111/medu.12773\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHatfield E, Cacioppo JT, Rapson RL (1993) Emotional Contagion. Curr Dir Psychol Sci 2:96\u0026ndash;100. https://doi.org/10.1111/1467-8721.ep10770953\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHe K, Zhang X, Ren S, Sun J (2016) Deep Residual Learning for Image Recognition\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHe Y, Tang TY (2017) Recommending highlights in Anime movies: Mining the real-time user comments \u0026ldquo;DanMaKu.\u0026rdquo; In: 2017 Intelligent Systems Conference (IntelliSys). pp 319\u0026ndash;322\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHorton D, Richard Wohl R (1956) Mass Communication and Para-Social Interaction. Psychiatry 19:215\u0026ndash;229. https://doi.org/10.1080/00332747.1956.11023049\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHossain G, Yeasin M (2017) Analysis of Cognitive Dissonance and Overload through Ability-Demand Gap Models. IEEE Trans Cogn Dev Syst 9:170\u0026ndash;182. https://doi.org/10.1109/TAMD.2015.2450681\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHu Y, Phawitpiriyakliti C, Terason S (2023) Tourist Satisfaction in Virtual Reality Immersive Experiences: Implications for the Tourism Industry. The Journal of Pacific Institute of Management Science (Humanities and Social Science) 9:115\u0026ndash;128\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHuang Z (2024) Study on the Optimization of Bilibili\u0026rsquo;s Business Model \u0026mdash;\u0026mdash; Based on the Business Model Canvas. HBEM 36:501\u0026ndash;508. https://doi.org/10.54097/z6hzst93\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eJun SH, Vogt C (2013) TRAVEL INFORMATION PROCESSING APPLYING A DUAL-PROCESS MODEL. Annals of Tourism Research 40:191\u0026ndash;212. https://doi.org/10.1016/j.annals.2012.09.001\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKAMEI K, TOYOTA A, KUSHIDA J (2012) Upsurge of Viewer\u0026rsquo;s Emotion by Video Sharing Using Pseudo Synchronization. J SOFT 24:944\u0026ndash;953. https://doi.org/10.3156/jsoft.24.944\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKan Y (2023) Research on the Integration of Multimodal Large Language Models (MLLM) and Augmented Reality (AR) for Smart Navigation with Real-Time Cross-Language Interaction and Cognitive Load Balancing Strategies. Journal of Business Research 5:2540752. https://doi.org/10.1142/S0129156425407521\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKim AJ, Ko E (2012) Do social media marketing activities enhance customer equity? An empirical study of luxury fashion brand. Journal of Business Research 65:1480\u0026ndash;1486. https://doi.org/10.1016/j.jbusres.2011.10.014\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLarson M, Melenhorst M, Men\u0026eacute;ndez M, Xu P (2014) Using crowdsourcing to capture complexity in human interpretations of multimedia content. In: Ionescu B, Benois-Pineau J, Piatrik T, Qu\u0026eacute;not G (eds) Fusion in Computer Vision: Understanding Complex Visual Content. Springer International Publishing, Cham, pp 229\u0026ndash;269\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLeng J, Zhu J, Wang X, Gu X (2016) Identifying the Potential of Danmaku Video from Eye Gaze Data. In: 2016 IEEE 16th International Conference on Advanced Learning Technologies (ICALT). pp 288\u0026ndash;292\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLeung D, Law R, van Hoof H, Buhalis D (2013) Social Media in Tourism and Hospitality: A Literature Review. Journal of Travel \u0026amp; Tourism Marketing 30:3\u0026ndash;22. https://doi.org/10.1080/10548408.2013.750919\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLi F, Wang W, Lai W, et al (2025) Unveiling the Multidimensional Nature of the Intention\u0026ndash;Behavior Gap. European Journal of Health Psychology 32:34\u0026ndash;50. https://doi.org/10.1027/2512-8442/a000162\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLi H, Tu X (2024a) Who generates your video ads? The matching effect of short-form video sources and destination types on visit intention. APJML 36:660\u0026ndash;677. https://doi.org/10.1108/APJML-04-2023-0300\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLi H, Tu X (2024b) Who generates your video ads? The matching effect of short-form video sources and destination types on visit intention. APJML 36:660\u0026ndash;677. https://doi.org/10.1108/apjml-04-2023-0300\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLi S, Scott N, Walters G (2015) Current and potential methods for measuring emotion in tourism experiences: a review. Current Issues in Tourism 18:805\u0026ndash;827. https://doi.org/10.1080/13683500.2014.975679\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLiu C (2024) A Study on the Correlation between Danmaku Videos and Audiences\u0026rsquo; Repetitive Viewing Behaviours - A Case Study of Bilibili Danmaku Video Network Videos. HC 1:. https://doi.org/10.61173/knqn8g59\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLiu Q, Xu L, Feng W, et al (2023) Is tourism live streaming a double-edged sword? The paradoxical impact of online flow experience on travel intentions. Journal of Travel \u0026amp; Tourism Marketing 40:744\u0026ndash;763. https://doi.org/10.1080/10548408.2023.2293016\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLiu W, Wang Z, Jian L, Sun Z (2024) How broadcasters\u0026rsquo; characteristics affect viewers\u0026rsquo; loyalty: the role of parasocial relationships\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLuo J (1986) The Influence of the Credibility of Brand Content on Social Media Platforms on Consumers\u0026rsquo; Purchasing Decisions and Its Communication Mechanism. Tourism Management 1:. https://doi.org/10.61173/9g99dd30\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLuo J, Chen M, Chen Y, et al (2021) Understanding and managing the threat of common method bias: Detection, prevention and control. Tourism Management 86:104330. https://doi.org/10.1016/j.tourman.2021.104330\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLyu B (2021) How is the Purchase Intention of Consumers Affected in the Environment of E-commerce Live Streaming?\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMalafouris L, Gosden C, Masson M, et al (2021) Building Social Media Engagement on Instagram by Using Visual Aesthetics and Message Orientation Strategy: A Content Analysis on Instagram Content of Indonesia Tourism Destinations. JICP 4:129\u0026ndash;138. https://doi.org/10.32535/jicp.v4i3.1304\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMattei E (2023) Multimodal corpus analysis of digital tourism narratives: A data-driven approach based on Systemic Functional Linguistics and social semiotics\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMehra P (2023) Unexpected surprise: Emotion analysis and aspect based sentiment analysis (ABSA) of user generated comments to study behavioral intentions of tourists. Tourism Management Perspectives 45:101063. https://doi.org/10.1016/j.tmp.2022.101063\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMehrabian A, Russell JA (1980) An Approach to Environmental Psychology. MIT Press, Cambridge, MA, USA\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMihaly Csikszentmihalyi (1990) Flow: The Psychology of Optimal Experience\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eNaab TK, Karnowski V, Schl\u0026uuml;tz D, et al (2019) Reporting Mobile Social Media Use: How Survey and Experience Sampling Measures Differ. Adaptive Behavior 13:126\u0026ndash;147. https://doi.org/10.1080/19312458.2018.1555799\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eNguyen PMB, Pham LX, Tran DK, Truong GNT (2024) A systematic literature review on travel planning through user-generated video. Journal of Vacation Marketing 30:553\u0026ndash;581. https://doi.org/10.1177/13567667231152935\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eOtto LP, Thomas F, Glogger I, De Vreese CH (2022) Linking Media Content and Survey Data in a Dynamic and Digital Media Environment \u0026ndash; Mobile Longitudinal Linkage Analysis. Digital Journalism 10:200\u0026ndash;215. https://doi.org/10.1080/21670811.2021.1890169\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePan X (2023) Motivations for game stream spectatorship: A content analysis of Danmaku on Bilibili. Global Media and China 8:190\u0026ndash;212. https://doi.org/10.1177/20594364231179750\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePascual-Fraile M del P, Villac\u0026eacute;-Molinero T, Tal\u0026oacute;n-Ballestero P, Chaperon S (2025) Post-pandemic collaborative destination marketing: Effectiveness and impact on different generational audiences. Front Hum Neurosci 31:665\u0026ndash;682. https://doi.org/10.1177/13567667231224091\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePeinado O, Shim M (2024) The intersection of \u0026ldquo;real\u0026rdquo; and \u0026ldquo;reel\u0026rdquo;: An investigation of K-pop idol dual self-presentation, paid advertisements, and fan engagement. Computers in Human Behavior 161:108414. https://doi.org/10.1016/j.chb.2024.108414\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePeinado O, Shim M, Alam A, et al (2020) Video Big Data Analytics in the Cloud: A Reference Architecture, Survey, Opportunities, and Open Research Issues. IEEE Access 8:152377\u0026ndash;152422. https://doi.org/10.1109/ACCESS.2020.3017135\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePeng G, Wang X, Li J, Wu J (2025) Is it beneficial for consumers to ask more questions in Danmaku? The inverted U-shaped effect of information-seeking Danmaku density on e-commerce livestream sales. Electron Commer Res. https://doi.org/10.1007/s10660-025-09959-1\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePertiwi E, Sanusi AP (2023) Storytelling in the Digital Age: Examining the Role and Effectiveness in Communication Strategies of Social Media Content Creators. Palakka Media Islam Commun 4:25\u0026ndash;34. https://doi.org/10.30863/palakka.v4i1.5082\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePham MT (2007) Emotion and Rationality: A Critical Review and Interpretation of Empirical Evidence. Review of General Psychology 11:155\u0026ndash;178. https://doi.org/10.1037/1089-2680.11.2.155\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePink A, Newton PM (2020) Decorative animations impair recall and are a source of extraneous cognitive load. Advances in Physiology Education 44:376\u0026ndash;382. https://doi.org/10.1152/advan.00102.2019\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePodsakoff PM, Organ DW (1986) Self-Reports in Organizational Research: Problems and Prospects. Journal of Management 12:531\u0026ndash;544. https://doi.org/10.1177/014920638601200408\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePolat E, \u0026Ccedil;elik F, Ibrahim B, K\u0026ouml;seoglu MA (2023) Unpacking the power of user-generated videos in hospitality and tourism: a systematic literature review and future direction. Journal of Travel \u0026amp; Tourism Marketing 40:894\u0026ndash;914. https://doi.org/10.1080/10548408.2023.2296655\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePop R-A, Săplăcan Z, Dabija D-C, Alt M-A (2022) The impact of social media influencers on travel decisions: the role of trust in consumer decision journey. Taylor \u0026amp; Francis\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRadford A, Kim JW, Hallacy C, et al (2003) Learning Transferable Visual Models From Natural Language Supervision\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRahwan I, Cebrian M, Obradovich N, et al (2024) Danmaku consistency reduces consumer purchases during live streaming: A dual-process model. Psychology and Marketing 41:2591\u0026ndash;2607. https://doi.org/10.1002/mar.22074\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRoss P, Paas F, Tuovinen JE, et al (2011) Is there an expertise of production? The case of new media producers. New Media \u0026amp; Society 13:912\u0026ndash;928. https://doi.org/10.1177/1461444810385393\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRoss P, Paas F, Tuovinen JE, et al (2003) Cognitive Load Measurement as a Means to Advance Cognitive Load Theory. Educational Psychologist 38:63\u0026ndash;71. https://doi.org/10.1207/S15326985EP3801_8\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSamson A, Voyer BG (2012) Two minds, three ways: dual system and dual process models in consumer psychology. AMS Rev 2:48\u0026ndash;71. https://doi.org/10.1007/s13162-012-0030-9\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSchwenzow J, Hartmann J, Schikowsky A, Heitmann M (2021) Understanding videos at scale: How to extract insights for business research. Journal of Business Research 123:367\u0026ndash;379. https://doi.org/10.1016/j.jbusres.2020.09.059\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eShah RR, Mahata D, Choudhary V, Bajpai R (2018) Multimodal Semantics and Affective Computing from Multimedia Content. Advances in Multimedia and Interactive Technologies 359\u0026ndash;382. https://doi.org/10.4018/978-1-5225-5246-8.ch014\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSigala M (2016) Social Media and the Co-creation of Tourism Experiences. The Handbook of Managing and Marketing Tourism Experiences 85\u0026ndash;111. https://doi.org/10.1108/978-1-78635-290-320161033\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSilaban PH, Chen W-K, Silaban BE, et al (2023) Demystifying Tourists\u0026rsquo; Intention to Visit Destination on Travel Vlogs: Findings from PLS-SEM and fsQCA. Emerg Sci J 7:867\u0026ndash;889. https://doi.org/10.28991/ESJ-2023-07-03-015\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSkulmowski A, Rey GD (2015) Measuring Cognitive Load in Embodied Learning Settings. Front Psychol 8:815\u0026ndash;827. https://doi.org/10.3389/fpsyg.2017.01191\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSpector PE (2019) Do Not Cross Me: Optimizing the Use of Cross-Sectional Designs. J Bus Psychol 34:125\u0026ndash;137. https://doi.org/10.1007/s10869-018-09613-8\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eStone AA, Shiffman S (2002) Capturing momentary, self-report data: A proposal for reporting guidelines. ann behav med 24:236\u0026ndash;243. https://doi.org/10.1207/S15324796ABM2403_09\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSU Y, LIU M, ZHAO N, et al (2021) Identifying psychological indexes based on social media data: A machine learning method. Adv Psychol Sci 29:571\u0026ndash;585. https://doi.org/10.3724/SP.J.1042.2021.00571\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSweller J (1988) Cognitive Load During Problem Solving: Effects on Learning. Cognitive Science 12:257\u0026ndash;285. https://doi.org/10.1207/s15516709cog1202_4\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTan-intaraarj P (2024) Exploring influence attempts, wishful identification, parasocial relationships, and behavioral loyalty among Thai game live-streamers and their viewers. Asian Journal of Communication 34:178\u0026ndash;194. https://doi.org/10.1080/01292986.2024.2315580\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eThein T, Westbrook RF, Harris JA (2022) How the associative strengths of stimuli combine in compound: Summation and overshadowing. Journal of Experimental Psychology: Animal Behavior Processes 34:155\u0026ndash;166. https://doi.org/10.1037/0097-7403.34.1.155\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTurner J, O\u0026rsquo;Leary M (2012) Targets\u0026rsquo; practices: how people allocate their attention among multiple streams of incoming information\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eUrry J (1990) The Tourist Gaze: Leisure and Travel in Contemporary Societies, 第 1st 版. SAGE Publications Ltd, London\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eVaswani A, Shazeer N, Parmar N, et al (2023) Attention Is All You Need\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWard L, Glancy M, Bowman S, Armstrong M (2020) The impact of new forms of media on production tools and practices\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWen H, Yang L, Estrin D (2019) Leveraging post-click feedback for content recommendations. In: Proceedings of the 13th ACM Conference on Recommender Systems. Association for Computing Machinery, New York, NY, USA, pp 278\u0026ndash;286\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWu G, Pei X, Wang D, et al (2024) I Bond, I Engage, I Visit: Investigating the Effects of Vloggers Tourist Engagement and Its Outcome on Tourist Attitudes. Journal of Travel Research. https://doi.org/10.1177/00472875241276546\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWu Q, Sang Y, Huang Y (2019) Danmaku: A New Paradigm of Social Interaction via Online Videos. Trans Soc Comput 2:1\u0026ndash;24. https://doi.org/10.1145/3329485\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWu S, Rizoiu M-A, Xie L (2018) Beyond Views: Measuring and Predicting Engagement in Online Videos. ICWSM 12:. https://doi.org/10.1609/icwsm.v12i1.15031\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWyer RS, Kardes FR (2020) A Multistage, Multiprocess Analysis of Consumer Judgment: A Selective Review and Conceptual Framework. J Consum Psychol 30:339\u0026ndash;364. https://doi.org/10.1002/jcpy.1158\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eXu P, Zhu X, Clifton DA (2023) Multimodal Learning With Transformers: A Survey. IEEE Trans Pattern Anal Mach Intell 45:12113\u0026ndash;12132. https://doi.org/10.1109/TPAMI.2023.3275156\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eYang J, Zeng Y, Liu X, Li Z (2022) Nudging interactive cocreation behaviors in live-streaming travel commerce: The visualization of real-time danmaku. Journal of Hospitality and Tourism Management 52:184\u0026ndash;197. https://doi.org/10.1016/j.jhtm.2022.06.015\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eYang Y (2020a) The \u003cem\u003edanmaku\u003c/em\u003e interface on Bilibili and the recontextualised translation practice: a semiotic technology perspective. Social Semiotics 30:254\u0026ndash;273. https://doi.org/10.1080/10350330.2019.1630962\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eYang Y (2020b) The danmaku interface on Bilibili and the recontextualised translation practice: a semiotic technology perspective. Taylor \u0026amp; Francis\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eYu Y, Huang S, Liu Y, Tan Y (2025) Emotions in Online Content Diffusion. Information Systems Research. https://doi.org/10.1287/isre.2022.0611\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eYuan H, Lu K, Ausaf A, Zhu M (2024) Constant or inconstant? The time-varying effect of danmaku on user engagement in online video platforms. INTR 35:771\u0026ndash;797. https://doi.org/10.1108/INTR-06-2023-0479\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eY\u0026uuml;ksel A (2017) A critique of \u0026ldquo;Response Bias\u0026rdquo; in the tourism, travel and hospitality research. Tourism Management 59:376\u0026ndash;384. https://doi.org/10.1016/j.tourman.2016.08.003\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZhou J, Zhou J, Ding Y, Wang H (2018) The Magic of Danmaku: A Social Interaction Perspective of Gift Sending on Live Streaming Platforms. SSRN Journal. http://dx.doi.org/10.2139/ssrn.3289119\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZhu S, Zhu X, Yao Y, Cheong CM (2025) Profiling the differences in strategy use in online multimodal reading: Associations with self-efficacy and reading task performance. Studies in Educational Evaluation 87:101507. https://doi.org/10.1016/j.stueduc.2025.101507\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"information-technology-and-tourism","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"jitt","sideBox":"Learn more about [Information Technology \u0026 Tourism](https://link.springer.com/journal/40558)","snPcode":"40558","submissionUrl":"https://submission.springernature.com/new-submission/40558/3","title":"Information Technology \u0026 Tourism","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"PUGC, Danmaku, travel intention, multimodal analysis, cognitive load, social interaction","lastPublishedDoi":"10.21203/rs.3.rs-8141450/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8141450/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eDigital tourism marketing increasingly relies on platform user-generated content (PUGC), yet mechanisms through which multimodal videos and real-time social interactions shape travel decisions remain largely opaque. This study investigates this \"black box\" by analyzing 2,650 tourism videos from Bilibili and their synchronized Danmaku (bullet-screen comments) through an innovative computational framework integrating computer vision, natural language processing, and machine learning. Three significant findings emerge: First, cognitive features demonstrate exceptional predictive power for travel intention (42.2% of model importance from only 10.4% of features), while traditional audiovisual features show limited contribution (16.1% importance from 45.8% of features)\u0026mdash;questioning visual-centric marketing assumptions. Second, non-linear models achieve substantially higher explanatory power than linear approaches (R\u0026sup2; = 0.518 vs. 0.033), revealing threshold effects and cognitive gating mechanisms not captured by additive frameworks. Third, Danmaku appears to function as cognitive scaffolding rather than distraction, with Granger causality suggesting orchestrated attention cascades from visual to audio to textual processing (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). By operationalizing previously unobservable mechanisms through extraction of 100\u0026thinsp;+\u0026thinsp;multimodal features aligned with 475\u0026nbsp;million Danmaku comments, this study provides new insights into platform affordances and tourism decision-making. The findings suggest important implications for digital tourism strategies: consider prioritizing cognitive activation alongside production aesthetics, facilitating synchronized social interaction, and recognizing that mental engagement may constitute a critical resource in attention economies.\u003c/p\u003e","manuscriptTitle":"From Watching to Wishing: A Multimodal Computational Analysis of How PUGC Video-Danmaku Ecology Shapes Travel Intention","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-12-04 07:21:42","doi":"10.21203/rs.3.rs-8141450/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-01-06T18:02:00+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-12-29T09:26:12+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-12-16T10:26:20+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"39497354291430942003914116697471195817","date":"2025-12-08T14:21:28+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"212759743394365103499222798387866433312","date":"2025-12-02T09:01:20+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-12-02T02:27:37+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-11-24T21:04:48+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-11-19T16:34:28+00:00","index":"","fulltext":""},{"type":"submitted","content":"Information Technology \u0026 Tourism","date":"2025-11-18T06:00:02+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"information-technology-and-tourism","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"jitt","sideBox":"Learn more about [Information Technology \u0026 Tourism](https://link.springer.com/journal/40558)","snPcode":"40558","submissionUrl":"https://submission.springernature.com/new-submission/40558/3","title":"Information Technology \u0026 Tourism","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"917d65c1-e12b-4c71-ba7e-49c12a652e20","owner":[],"postedDate":"December 4th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-05-13T12:11:44+00:00","versionOfRecord":[],"versionCreatedAt":"2025-12-04 07:21:42","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8141450","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8141450","identity":"rs-8141450","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.