Cultural Faithfulness in Tourism Chatbots: A Structured Human Adjudication Framework for Traceable Retrieval-Augmented Generation | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Cultural Faithfulness in Tourism Chatbots: A Structured Human Adjudication Framework for Traceable Retrieval-Augmented Generation Hindun Nurhidayati -, Zulin Nurchayati -, Sugiarto - This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8994226/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract The integration of large language model-based chatbots into tourism contexts has sparked critical discussions about Cultural Faithfulness, especially concerning the accurate representation of intangible heritage. While Retrieval-Augmented Generation (RAG) enhances factual accuracy, current evaluation methods predominantly depend on automated similarity metrics and seldom incorporate structured human adjudication, often conflating semantic coherence with epistemic validity. To address these limitations, this study proposes a Human-in-the-Loop evaluation framework for traceable RAG systems in tourism chatbots. The framework combines chunk-level retrieval traceability, a granular narrative error taxonomy (E0–E10) designed to capture varying degrees of attribution erosion, automated similarity filtering, and human-based documented consensus adjudication into a cohesive protocol. By treating retrieval as an epistemic constraint on generative processes and operationalizing abstention as a measure of boundary awareness, the framework establishes rigorous evaluation criteria. Empirical validation was conducted using 61 structured questions derived from a corpus of Indonesian cultural narratives, generating 183 independent annotations. Analysis revealed that 73.8% of responses met acceptance criteria after artifact-level review, while 26.2% were excluded due to grounding violations. Notably, high-severity errors were confined to rejected cases, and iterative testing confirmed no progression into hallucination or contradiction categories. These results confirm that Cultural Faithfulness can be systematically achieved through traceable retrieval mechanisms, structured human validation, and governance-aligned artifact preservation. This research extends evaluation methodologies beyond superficial similarity metrics, advancing a unified model of epistemic accountability for generative systems in tourism applications. Cultural Faithfulness Artifact-Level Traceability Attribution Erosion Structured Human Adjudication Narrative Boundary Stability Epistemic Accountability Figures Figure 1 1. Introduction The digital transformation in tourism has driven the adoption of large language model (LLM) chatbots, which now serve dual roles as information providers and narrative agents. These systems influence how tourists engage with folklore, local history, and traditional practices, raising critical questions about their cultural responsibilities. When a chatbot represents a culture, factual accuracy is insufficient; cultural communication requires fidelity to source narratives, contextual sensitivity, and epistemic accountability for knowledge representation. International frameworks emphasize the protection of intangible cultural heritage. The 2003 UNESCO Convention identifies oral traditions as vital domains requiring preservation through intergenerational transmission and protection from distortion (UNESCO, 2003 ). In artificial intelligence contexts, this raises a contemporary challenge: how to ensure digital agents preserve cultural values when communicating folklore? UNESCO ( 2025 ) warns that algorithms cannot fully replace human values, underscoring the need for cultural guardians to retain control over heritage representation. Retrieval-Augmented Generation (RAG) architectures aim to improve generative models' factual precision by conditioning outputs on external documents (Lewis et al., 2020 ; Gao et al., 2024 ). However, evaluation remains dominated by automated metrics like ROUGE (Lin, 2004 ) and BERTScore (Zhang et al., 2020 ), which assess surface-level similarity rather than epistemic alignment with source evidence. Cross-cultural NLP research reveals that large language models may produce grammatically coherent outputs misaligned with local community values, highlighting cultural gaps in current AI systems (Cao et al., 2023 ). The limitations inherent in digital storytelling systems carry significant implications for cultural tourism. Subtle interpretive expansions—such as the inclusion of unverified details, oversimplification of narrative arcs, or moral reframing—pose a risk of displacing narrative authority from cultural communities to algorithmic systems. This phenomenon becomes particularly critical when chatbots function as digital narrators within culturally rich environments, where insufficient accountability mechanisms could fundamentally alter knowledge transmission dynamics (Nurhidayati et al., 2025 ). The central issue transcends basic attribution erosion detection, revealing a systemic lack of diagnostic frameworks for multi-layered grounding failures. These failures manifest as attribution erosion, unsubstantiated expansions, or outright factual contradictions. While prior research has addressed hallucination detection (Ji et al., 2023 ) and the impact of inaccurate generative outputs on consumer trust in tourism (Kim et al., 2025 ), three critical structural deficiencies persist: First, current evaluation protocols rarely combine artifact-level traceability with systematic human oversight. Automated evaluation tools may identify semantic inconsistencies but remain ineffective at detecting attribution errors, unsubstantiated expansions, or boundary shifts in narrative frameworks. Second, Human-in-the-Loop (HITL) implementations predominantly focus on model optimization rather than epistemic validation. In conventional systems, human review operates as an after-the-fact verification process rather than a structured evaluation phase with defined diagnostic criteria. Contemporary critiques highlight that mere human involvement does not ensure ethical or accurate system behavior; this requires explicit role delineation and comprehensive information support for adjudicators (Tschiatschek et al., 2024 ). Third, governance principles like transparency and accountability are frequently stated in abstract terms without concrete implementation through artifact-level traceability or documented consensus adjudication protocols. The 2024 OECD AI Principles explicitly recognize traceability as an accountability mechanism for outcome analysis and inquiry resolution (OECD, 2024 ). Despite this, human-mediated adjudication—intended to operationalize these principles—remains underdeveloped in Retrieval-Augmented Generation (RAG) systems. The proposed evaluation framework for Retrieval-Augmented Generation (RAG) in tourism chatbots addresses critical gaps by integrating four methodological innovations. First, source traceability is preserved through artifact-level chunk injection, where cultural narratives are explicitly embedded into the model context. This transforms retrieval from probabilistic similarity matching to deterministic evidence provision, a foundational requirement for epistemic accountability (Hickey, 2026 ). Second, a granular error classification system (E0–E10) was developed to assess grounding severity. The taxonomy differentiates between fully supported responses (E0), minor omissions (E1), partial grounding with untraceable additions (E2), and severe misrepresentations like contradictions (E8) or attribution erosion (E7). This replaces binary true/false evaluations with a diagnostic framework that captures nuanced error types. Third, automated metrics such as BERTScore and ROUGE-L function as analytical supports rather than final judgments. These generate diagnostic artifacts that inform human evaluators, addressing critiques about the misalignment between automated metrics and human assessments of faithfulness (Maynez et al., 2020 ). Finally, epistemic validation is formalized through documented consensus adjudication. Annotators collectively decide on acceptance, rejection, or abstention using co-reasoning (Salloch & Eriksen, 2024 ). The abstention mechanism explicitly signals boundary awareness, aligning with symbiotic epistemology (Kapusta, 2025 ) and ethical obligations to avoid cultural misrepresentation (Floridi et al., 2018 ). The proposed evaluation framework transcends conventional output scoring by functioning as a systematic methodology to assess whether retrieval mechanisms impose epistemic constraints on generative processes. Cultural Faithfulness is implemented through three core components: components: artifact-level traceability for source verification, severity-differentiating diagnostics for error classification, and reproducible artifact preservation for accountability. To ensure methodological rigor, the evaluation dataset has been archived as an open research artifact (see Supplementary Material A), containing structured queries, granular annotation logs, chunk-level mapping data, and consensus metadata. This study investigates five central research questions (RQ): What is the discriminative capacity of the evaluation framework in distinguishing grounded responses from attribution erosion and severe grounding violations? How does automated similarity filtering complement structured human adjudication rather than supplant it? In what ways does the E0–E10 taxonomy function as a diagnostic tool for assessing grounding quality? What role does chunk-injection retrieval play in maintaining narrative boundary stability during iterative generation processes? How does artifact preservation operationalize transparency and accountability in culturally situated chatbot systems? This study presents four key contributions to the fields of smart tourism and responsible artificial intelligence. First, it expands the understanding of Cultural Faithfulness as an evaluative framework specifically tailored for tourism chatbots. By redefining chatbots as culturally embedded narrative facilitators rather than mere information providers, the research connects digital tourism scholarship with cross-cultural natural language processing (NLP) assessment methodologies. Second, the study proposes a spectrum-based diagnostic classification system (E0–E10) to evaluate attribution erosion in retrieval-augmented generation (RAG) systems. This taxonomy moves beyond simplistic true/false accuracy metrics by systematically categorizing response types including grounded assertions, contextual transitions, alignment errors, attribution erosion, contradictory statements, deliberate omissions, and systemic malfunctions. Third, the research implements narrative boundary stability through a structured human adjudication architecture incorporating chunk injection. This approach ensures artifact-level traceability by directly linking generated content to source text segments while maintaining evaluation records. The methodology offers a replicable template for assessing culturally sensitive generative systems. Fourth, the framework advances tourism AI governance by converting abstract transparency and accountability principles into actionable evaluation protocols. Through version-controlled dataset management, documented consensus adjudication practices, and annotation-level record preservation, the study demonstrates practical governance implementation in tourism AI applications. These contributions collectively forge an interdisciplinary nexus between smart tourism research, retrieval-augmented language modeling, structured human evaluation, and operational AI governance. 2. Literature Review and Theoretical Background 2.1 Digital Tourism, Cultural Mediation, and Epistemic Accountability Digital transformation has permeated nearly all sectors, with tourism communication undergoing significant reconfiguration. Smart systems now mediate interactions between destinations and travelers, transforming tourism platforms from passive information repositories into culturally rooted narrative agents. Gretzel et al. ( 2015 ) documented how this evolution reconfigures communication paradigms. Smart tourism ecosystems have institutionalized data-driven interactivity, algorithmic personalization, and the digital co-creation of experiences (Sigala, 2018 ). Recent studies emphasize the integration of AI into destination management practices, including chatbots that operationalize transparency and accountability (Caric et al., 2026 ) and culturally embedded facilitators within smart tourism infrastructures (Cruz et al., 2025 ). These developments demonstrate that algorithmic mediation has transitioned from peripheral to structurally embedded within tourism communication systems. Despite these advancements, smart tourism research predominantly emphasizes competitiveness, operational efficiency, personalization, and sustainability. Consequently, systematic attention to what can be termed "epistemic accountability" —the responsibility of a knowledge system to accurately represent its sources and the boundaries of its own knowledge —within cultural representation remains relatively limited. This gap is particularly consequential given the normative obligations attached to the protection of intangible cultural heritage, which demand authenticity, contextual coherence, and faithful community representation (UNESCO, 2003 ). When generative systems, such as tourism chatbots, construct cultural narratives, they inevitably assume an interpretive authority that extends beyond mere factual transmission. They make choices about what to include, exclude, and emphasize, thereby shaping the user's understanding of the culture. In the domain of digital tourism, intelligent systems function as critical mediators that fundamentally alter how travelers perceive and interact with cultural heritage (Gretzel, 2011 ). From the perspective of human-centered AI and responsible innovation, such systems must be evaluated not only by their performance metrics but also by their alignment with broader social and ethical expectations (Shneiderman, 2020; Floridi et al., 2018 ). Beyond mere technical performance, the development of culturally faithful chatbots must be grounded in the Responsible Innovation (RI) framework, which advocates for a proactive and reflective approach to technological advancement (Stilgoe et al., 2013 ). This evaluation must scrutinize not just whether the information is factually correct, but whether the system's mode of presentation and its handling of uncertainty respect the epistemic integrity of the source culture. Consequently, traceability becomes a primary ethical requirement, ensuring that every cultural claim can be audited against authoritative source evidence. By applying the four pillars of RI—anticipation, reflexivity, inclusion, and responsiveness—the proposed framework moves beyond static evaluation. It ensures that the system's retrieval mechanisms are not only transparent but also responsive to the evolving ethical and social expectations of the source communities, thereby safeguarding the epistemic integrity of cultural knowledge through structured human oversight. As Tsamados et al. ( 2022 ) highlight in their updated analysis of algorithmic ethics, alongside well-documented concerns about bias and fairness, issues of "epistemic concerns"—including the justification of knowledge claims and the opacity of decision-making—are central to the trustworthiness of AI systems (see also Mittelstadt et al., 2016 ). Therefore, a framework for cultural chatbots must go beyond conventional evaluation to incorporate mechanisms for traceability and structured human oversight, ensuring that the system's epistemic authority is both accountable and aligned with the values of the communities whose heritage it mediates. This shift raises questions concerning epistemic accountability in algorithmic mediation rather than just factual accuracy. Cross-cultural communication theory emphasizes that the construction of meaning is mediated by culturally rooted interpretive frameworks (Ting-Toomey, 2015 ). Cultural dimensions theory operationalizes variations in value orientations, power distance, and uncertainty tolerance across societies (Hofstede et al., 2010 ). In natural language processing research, cultural alignment is increasingly distinguished from linguistic fluency, necessitating context-sensitive adaptation mechanisms (Hershcovich et al., 2022 ). Empirical studies have demonstrated measurable gaps between the outputs of generative models and the value distributions of cross-cultural societies (Cao et al., 2023 ; Lee et al., 2025 ). These findings suggest that generative systems may remain grammatically coherent while being epistemically misaligned with local cultural norms. Consequently, tourism chatbots should not be evaluated solely as information delivery tools. They must be understood as culturally rooted speakers operating within normative and governance constraints. While the smart tourism literature acknowledges digital mediation, systematic mechanisms for assessing Cultural Faithfulness in generative tourism systems remain underdeveloped, particularly at the intersection of epistemic accountability and cultural representation. Considering this cultural context, we now examine the technical framework that facilitates these narratives through culturally rooted narrative agents and epistemically constrained generative processes. 2.2 Generative Models, Attribution Erosion, and Epistemically Constrained Retrieval-Augmented Generation This section examines the technical characteristics of generative models within the cultural context established earlier. While these systems demonstrate significant creative capacity, they remain vulnerable to attribution erosion, unsubstantiated elaborations, and groundless fabrications (Ji et al., 2023 ; Bender et al., 2021 ). These risks pose critical challenges in domains where narrative authority is culturally codified. Beyond surface-level inaccuracies, such models may generate seemingly coherent explanations that lack verifiable evidential foundations, reflecting a prioritization of performance over epistemic fidelity (Birhane et al., 2022 ). Retrieval-Augmented Generation (RAG) mitigates these limitations through epistemic constraints that bind the generative process to external knowledge sources By anchoring outputs in retrieved evidence (Lewis et al., 2020 ; Guu et al., 2020 ), subsequent improvements have refined retrieval conditioning, optimized knowledge integration, and enhanced scaling efficiency (Borgeaud et al., 2022 ). Multi-hop architectures now enable contextual aggregation across distributed knowledge repositories (Asai et al., 2022 ), while artifact-level traceability systems establish explicit links between generated content and source evidence segments (Tafjord et al., 2021 ). Explainable RAG frameworks further advance this approach by documenting intermediate reasoning pathways and retrieval decisions, thereby supporting auditability and transparency (Ren et al., 2025 ). These methodological advances collectively reinforce narrative boundary stability and reduce ungrounded generation risks. However, factual grounding does not inherently ensure epistemic accountability. Cross-cultural NLP research demonstrates that models may retrieve factually accurate information while subtly reframing narratives in culturally misaligned or normatively divergent ways (Hershcovich et al., 2022 ; Liu et al., 2025 ). Surveys of evaluation frameworks reveal that benchmark-driven RAG assessments often prioritize lexical overlap and answer matching, frequently compromising contextual fidelity and narrative boundary stability (Gao et al., 2024 ). Studies on factuality in summarization show that surface-level alignment with source texts does not guarantee cultural faithfulness, which systematically distinguishes between literal factual adherence and the preservation of culturally codified meaning (Maynez et al., 2020 ). Standard automated metrics such as ROUGE (Lin, 2004 ) and BERTScore (Zhang et al., 2020 ) measure semantic similarity but fail to detect plausible yet unsubstantiated reasoning, epistemic shifts, or culturally sensitive narrative distortions. Research on attribution erosion typologies indicates that generative errors exist on a spectrum from minor contextual deviations to severe fabrications, including epistemic misrepresentations that similarity-based metrics often overlook (Ji et al., 2023 ). Furthermore, bias amplification and representational harms complicate evaluation, especially in culturally sensitive domains where narrative authority is culturally codified (Abid et al., 2021 ). These distortions often arise when generative models fail to preserve the specific cultural and temporal nuances of word meanings, leading to narrative misalignment (Lucy & Bamman, 2021 ). Retrieval-Augmented Generation (RAG) mitigates these limitations through epistemic constraints. Recent studies highlight that improved retrieval quality does not inherently produce epistemically constrained generative outputs. Salemi and Zamani ( 2024 ) demonstrate that systematic evaluation of retrieval components is essential for anchoring outputs in retrieved evidence within retrieval-augmented generation (RAG) systems. However, high retrieval performance does not ensure that generated narratives maintain contextual anchoring or cultural codification. Even with consistent retrieval accuracy, epistemically constrained generative processes may introduce interpretive shifts, subtle reframing, or contextual expansions. These findings emphasize the need for structured evaluation frameworks that assess both artifact-level traceability and the stability of cultural codification in generated narratives. While RAG significantly enhances artifact-level traceability, current evaluation paradigms remain inadequate for detecting nuanced attribution erosion, culturally rooted reframing, and longitudinal shifts in narrative boundary stability. 2.3 Human Evaluation, Reliability, and Spectrum-Based Diagnosis Automated metrics remain insufficient for capturing contextual nuances in culturally mediated AI systems, necessitating human evaluation as a critical validation mechanism. Inter-rater agreement metrics such as Cohen's kappa (Cohen, 1960 ) and Fleiss' kappa (Fleiss, 1971 ) operationalize annotation consistency across evaluators, while computational linguistics research formalizes this annotation reliability within NLP evaluation pipelines (Artstein & Poesio, 2008 ). Interpretive benchmarks (Landis & Koch, 1977 ; Krippendorff, 2019 ) establish structured thresholds for evaluating agreement strength, but in complex qualitative assessments involving extensive taxonomies, this rigor is further bolstered through negotiated agreement and intercoder reliability processes (Campbell et al., 2013 ). These methodological foundations collectively ensure the systematic integrity required for evaluating complex categorical taxonomies in AI-generated content. The structured human adjudication paradigm integrates documented consensus adjudication into AI workflows, enabling iterative correction and supervision (Amershi et al., 2019 ). From a human-centered AI perspective, oversight mechanisms must prioritize both system performance optimization and the enforcement of accountability through controlled deployment (Shneiderman, 2020). However, current implementations often emphasize model refinement and usability improvements over the systematic adjudication of narrative epistemic accountability in culturally sensitive domains. Critiques of conventional human-in-the-loop models highlight the inadequacy of assuming human presence as a default guarantee for correct and ethical system function (Tschiatschek et al., 2024 ). A strategic transition is required, where the human role as a decision-maker is explicitly defined and supported by comprehensive information. This aligns with the concept of co-reasoning proposed by Salloch and Eriksen ( 2024 ), wherein experts engage in collaborative evidence interpretation to achieve legitimate epistemic decisions through documented consensus. Studies on dialectal and sociolinguistic variations demonstrate how subtle linguistic deviations can encode cultural distortions or marginalization (Blodgett et al., 2020 ). Complementary research on attribution erosion typologies and epistemically constrained retrieval-augmented generation (RAG) indicates that generative errors are neither binary nor uniform; instead, they are distributed along a gradational continuum of attribution erosion (Ji et al., 2023 ). The comprehensive survey by Ji et al. ( 2023 ) systematically catalogues the landscape of hallucination phenomena in natural language generation, providing a foundational framework for understanding how generative errors manifest across tasks and why evaluation must extend beyond surface-level metrics. These findings suggest that binary truth labels fail to capture the complexity of narrative misalignment. Collectively, this research trajectory justifies a spectrum-based diagnosis of attribution erosion. Rather than classifying outputs as simply true or false, a gradational scheme allows for the systematic differentiation between grounded responses, minor contextual shifts, moderate cultural misalignment, and severe fabrication. Such spectrum-based operationalization aligns with annotation reliability theory, bias taxonomy design, and structured human adjudication paradigms. Consequently, this approach provides the theoretical foundation for a multi-level diagnostic evaluation framework in culturally mediated generative systems, where narrative boundary stability is preserved through artifact-level traceability and epistemic accountability. 2.4 Narrative Boundary Stability and Artifact-Level Traceability Artifact-level traceability systems enhance evidential transparency, yet the structural stability of narrative boundaries during epistemically constrained generative processes remains under-theorized. Generative systems conditioned on retrieved chunks frequently extrapolate beyond evidential limits, creating a hybridization of grounded content and unsubstantiated elaborations. This dynamic reflects an inherent conflict between probabilistic language modeling and epistemic constraints. Retrieval-augmented architectures improve contextual anchoring through external knowledge integration during generation (Guu et al., 2020 ; Lewis et al., 2020 ). Large-scale fine-tuning subsequently optimizes retrieval depth and parameter efficiency (Borgeaud et al., 2022 ). Multi-hop reasoning systems further enhance cross-document aggregation and contextual linkage (Asai et al., 2022 ). Nevertheless, increased retrieval capacity does not inherently prevent generative extrapolation beyond epistemic boundaries. Studies of plausible reasoning errors show that models often produce semantically fluent yet unsubstantiated expansions (Ji et al., 2023 ). This highlights a crucial distinction between artifact-level traceability and narrative boundary stability. Even with relevant artifacts retrieved, systems may generate extensions that subtly transgress evidential constraints while maintaining rhetorical coherence. The concept of boundary awareness, introduced within the symbiotic epistemology framework (Kapusta, 2025 ), offers a means to comprehend this phenomenon. Kapusta posits that a reliable epistemic agent is one capable of recognizing the limits of its knowledge and refraining from action when evidence is insufficient. The dual-level transparency principle in his TRACE protocol, encompassing high-level reasoning patterns and detailed factor explanations, provides a language for describing how boundary awareness can be operationalized. Furthermore, Hickey ( 2026 ) argues that epistemic trust in opaque AI systems can be fostered through the provision of authoritative and verifiable artifact-level traceability systems. This approach diverges from efforts to demystify internal algorithmic black boxes. Within the RAG context, retrieved chunks serve as these constitutive knowledge sources. Technical explorations of chunk injection mechanisms for cultural tourism applications have demonstrated the feasibility of this approach, though systematic evaluation frameworks remain underdeveloped. By explicitly embedding chunks as contextual memory, artifact-level traceability becomes the bedrock for epistemic accountability. Narrative boundary stability refers to the degree to which generated narratives remain confined within the epistemic scope defined by retrieved artifacts. Artifact-level traceability systems enhance traceability by explicitly linking outputs to ranges of supporting evidence (Tafjord et al., 2021 ). However, existing evaluation frameworks seldom assess how consistently generative outputs respect artifact-level traceability at the chunk level. By conceptualizing chunk injection as a contextual anchoring mechanism, narrative boundary stability emerges as an operationalizable construct linking artifact-level traceability with narrative containment. This framing extends the grounding literature by introducing artifact-level diagnostics that evaluate not only whether information is retrieved, but whether generative extrapolations remain epistemically constrained within the retrieved scope. Such artifact-sensitive evaluations advance the theoretical integration of retrieval conditioning and narrative governance in generative systems. 2.5 Governance, Cultural Faithfulness, and Epistemic Accountability The AI governance framework operationalizes transparency, accountability, fairness, and meaningful human control as foundational principles for responsible deployment (OECD, 2019 ; Jobin et al., 2019 ). The artifact-level traceability principle in the OECD AI Principles 2024 explicitly positions traceability as an accountability instrument to enable the analysis of outcomes and responses to inquiries (OECD, 2024 ). AI ethics scholarship further argues that governance frameworks must extend beyond abstract ethical principles to incorporate culturally codified epistemic accountability, particularly in contexts where knowledge production intersects with identity and heritage (Mohamed, Png, & Isaac, 2020 ). Within tourism mediation settings, tensions between algorithmic storytelling and indigenous interpretive authority have been empirically documented. Previous studies examining automated systems in cultural contexts have demonstrated how such systems can inadvertently reshape cultural representations through generative processes (Nurhidayati et al., 2025 ). However, technical efforts to enhance retrieval reliability in tourism chatbots suggest that epistemic accountability is operationally feasible. A broader epistemic analysis of generative tourism systems reveals persistent risks of narrative compression, oversimplification, and cultural distortion despite retrieval augmentation (Nurhidayati et al., 2026 ). These findings indicate that governance frameworks must extend beyond high-level ethical principles to implement structured human adjudication paradigms that preserve narrative boundary stability through artifact-level traceability systems. Operational epistemic accountability necessitates the operationalization of governance principles into reproducible evaluative procedures. This requires severity-differentiating diagnostics, structured human adjudication protocols, and reproducible artifact preservation to enable cross-version comparison and longitudinal auditability. The principles of reproducibility and artifact-level traceability in computational research increasingly underscore the need to document methodological evolution and maintain evidential outputs for verification. In culturally sensitive domains, governance becomes inseparable from evaluation architecture. Artifact-level traceability systems, severity-differentiating diagnostics for epistemic deviations, and preserved diagnostic artifacts constitute concrete instantiations of governance principles. The AI4People framework proposed by Floridi and colleagues ( 2018 ) provides an ethical foundation with the principles of beneficence, non-maleficence, autonomy, justice, and epistemic accountability. The principle of epistemic accountability, in particular, demands that AI systems be explainable and answerable. This demand is met by the artifact-level traceability systems within our framework. Furthermore, Laux and colleagues' ( 2024 ) critique of equating trustworthiness with risk acceptability in the EU AI Act highlights the need for a richer definition of trust in AI. Our framework contributes to this discourse by operationalizing Cultural Faithfulness as a crucial component of trustworthiness within the context of cultural tourism. The Human-Centered AI vision by Shneiderman ( 2022 ) emphasizes the importance of human control and transparent reporting. This vision is embodied in our framework's design. The principles of artifact-level traceability and structured human adjudication that we implement are concrete manifestations of that vision for cultural domains. By embedding governance into evaluation design, cultural protection shifts from declarative policy aspirations to operational accountability. 2.6 Research Gaps and Integrative Positioning The reviewed literature reveals fragmentation across four interconnected domains: digital tourism communication, retrieval engineering, human evaluation methodologies, and AI governance. Smart tourism research acknowledges AI-mediated storytelling and digital interactivity (Gretzel et al., 2015 ; Sigala, 2018 ). However, the systematic operationalization of Cultural Faithfulness as an evaluative construct remains limited. Retrieval-Augmented Generation (RAG) architectures enhance factual grounding through contextual anchoring (Lewis et al., 2020 ; Gao et al., 2024 ). Yet, evaluation protocols predominantly emphasize lexical similarity and benchmark optimization (Lin, 2004 ; Zhang et al., 2020 ). Reliability theory formalizes inter-rater consensus (Artstein & Poesio, 2008 ), although structured epistemic diagnostics are seldom integrated into the evaluation of applied generative systems. Concurrently, governance frameworks articulate principles of transparency and accountability (OECD, 2019 ; Jobin et al., 2019 ). These norms are rarely operationalized through artifact-level evaluation protocols. A critical limitation in current scholarship lies in the absence of a unified evaluative architecture that systematically integrates artifact-level traceability systems, severity-differentiating diagnostics, structured human adjudication protocols, and reproducible artifact preservation mechanisms. While prior research has examined these elements in isolation, their coordinated implementation remains insufficiently developed, especially in culturally sensitive domains such as tourism. This research proposes a traceable structured human adjudication framework for epistemically constrained generative processes in tourism chatbots. The framework combines chunk injection as contextual anchoring mechanisms, severity-differentiating diagnostics, consensus-based reliability validation, and artifact-level traceability systems. By operationalizing Cultural Faithfulness through epistemic accountability principles, this work establishes an interdisciplinary nexus between smart tourism research and responsible AI evaluation, ensuring narrative boundary stability through artifact-level traceability and epistemic constraint enforcement. 3. Methodology 3.1 Research Design and Analytical Scope This study employs a structured evaluative design grounded in the structured human adjudication paradigm. Its primary objective extends beyond measuring semantic similarity to examining whether retrieval functions as epistemic constraints on epistemically constrained generative processes within culturally codified tourism narratives. The analysis unit comprises 61 system responses corresponding to 61 evaluation questions. Each response undergoes two complementary analytical procedures: automated similarity screening and structured human adjudication using severity-differentiating diagnostics. Two levels of analysis are defined. At the question level, 61 final consensus decisions are examined to determine grounding classification patterns. At the annotation level, 183 independent scores generated by three evaluators are analyzed to assess diagnostic distribution and agreement structure. This dual structure ensures the evaluation captures both output categorization and more granular epistemic deviation patterns. 3.2 Dataset Construction and Question Design The evaluation dataset (Version 1.2) was constructed from curated cultural documents sourced from the Ande-Ande Lumut narrative (see Supplementary Material A for the complete dataset). As discussed in the Theoretical Foundation, Cultural Faithfulness in RAG systems requires epistemically constrained references. Consequently, the narrative corpus functions as a controlled ground-truth knowledge base rather than a general domain benchmark. Prior to retrieval indexing, source documents were segmented into semantically coherent text units. Each segment was recorded in the chunk segmentation table containing unique identifiers, exact text ranges, and their positions within the source documents. This segmentation enables artifact-level traceability and direct mapping between generated claims and evidential ranges. Table 1 Distribution of Question Types in Dataset v1.2 (n = 61) Question Type Number Analytical Purpose Related RQ In-context 30 Measure factual grounding accuracy and evidence alignment RQ1 Ambiguous 16 Test contextual sensitivity and interpretative restraint RQ1 Out-of-context 15 Test attribution erosion and boundary control RQ1 Total 61 Structured diagnostic evaluation dataset - Table source: Research data (can be verified in Supplementary Material A) As shown in Table 1 , the dataset comprises 30 in-context questions, 16 ambiguous questions, and 15 out-of-context questions. In-context questions assess the accuracy of grounding in the available evidence. Ambiguous questions evaluate contextual sensitivity and interpretative restraint. Out-of-context questions test attribution erosion behavior and resistance to unsubstantiated expansions. These three categories operationalize the attribution erosion typologies identified in the Theoretical Foundation. To test the stability of narrative boundaries under repeated generations, a subset of 14 questions was administered again under controlled conditions. Table 2 Distribution of Repeated Questions (n = 14) Question Category Number Repeated Analytical Purpose Related RQ In-context 7 Test grounding stability under repetition RQ4 Ambiguous 4 Test interpretative consistency RQ4 Out-of-context 3 Assess abstention mechanism reliability RQ4 Total Repeated 14 Narrative boundary stability analysis - Table source: Research data (can be verified in Supplementary Material A) Repeated questions enable the assessment of artifact-level traceability and epistemic constraint enforcement across diverse generations. 3.3 Retrieval Architecture with Chunk Injection and Structured Human Adjudication . The proposed evaluation architecture integrates retrieval, generation, filtering, and structured human adjudication into a unified workflow, as illustrated in Fig. 1. [ Insert Fig. 1 here ] Figure 1 . Structured Human Adjudication Architecture for Cultural Faithfulness, Showing Retrieval-Augmented Generation (RAG) Components The proposed framework integrates four core methodological components: (1) artifact-level traceability through chunk injection, (2) severity-differentiating diagnostics using the E0–E10 taxonomy, (3) structured human adjudication with documented consensus procedures, and (4) reproducible artifact preservation. The architecture operates as follows. Culturally codified knowledge sources ① are segmented into indexed text chunks ② and explicitly injected into model context ③, establishing artifact-level traceability. A single query triggers two parallel retrieval-augmented generation (RAG) paths. The RAG system (Candidate Generator) ④ produces candidate responses, while the RAG system (Reference & Chunk_ID Generator) ⑤ generates reference responses accompanied by supporting chunk identifiers. This dual-path configuration enables direct mapping between generated claims and source evidence. An automated metric filtering layer ⑥ computes similarity scores, including BERTScore F1, ROUGE-L F1, and cosine similarity, between candidate and reference responses. These metrics function solely as analytical signals and do not incorporate chunk identifiers, thereby avoiding their use as grounding validators. Threshold-based similarity labeling provides early indicators of potential epistemic deviation. In the dataset, these values are recorded as bertscore_f1, rougeL_f1, and cosine_similarity. Subsequently, three evaluators independently perform structured human adjudication ⑦ using the E0–E10 severity-differentiated taxonomy. Evaluators are granted full access to source documents, chunk segmentation tables, retrieved chunk identifiers, candidate and reference responses, and automated filtering outputs. Cases meeting predefined trigger conditions, such as evaluator divergence or high-severity classifications, undergo documented consensus adjudication ⑧ to achieve unanimous final decisions. Individual annotations are stored with evaluator_id, error_code, and annotation_notes fields. All artifacts generated throughout the pipeline, including source documents, segmentation tables, similarity metrics, and consensus resolutions, are systematically archived under reproducible artifact preservation ⑨. The archive includes 183 independent annotations, 61 final decisions, evaluation questions, generated responses, chunk mappings, filtering outputs, and consensus metadata. The complete anonymized archive is provided as Supplementary Material A. This architecture operationalizes narrative boundary stability by verifying whether generated outputs remain confined within epistemically constrained evidential ranges during generation. The structured adjudication layer embeds interpretative validation directly within the retrieval workflow, addressing the evaluative integration gap identified in the Theoretical Foundation. 3.4 Automated Similarity Filtering Post-generation, candidate responses and reference responses underwent automated similarity analysis. Similarity metrics were computed using BERTScore F1, ROUGE-L F1, and cosine similarity. Threshold-based similarity classifications and early indicators of potential grounding issues were generated to support interpretive analysis. In the dataset, these are recorded as similarity_label (high/medium/low) and early_warning_flags. Notably, similarity calculations were restricted to pairwise comparisons between candidate and reference responses. Chunk identifiers were excluded from metric computations, ensuring that filtering functioned as an analytical signal rather than a definitive grounding validator. Automated metrics did not dictate final evaluation outcomes. Instead, these metrics provided structured input for documented consensus adjudication. All metric outputs were preserved within the evaluation archive (Supplementary Material A). 3.5 Granular Cultural Faithfulness Scale (E0–E10) Narrative boundary stability is evaluated through an 11-point diagnostic framework spanning E0 (fully grounded) to E10 (system failure). This severity-differentiating taxonomy surpasses binary detection approaches by identifying nuanced degrees of epistemic constraint violation and unsubstantiated expansions. Table 3 Granular Cultural Faithfulness Scale (E0–E10) Code Category Description Epistemic Severity Failure Type E0 Fully Grounded Response is fully relevant, correctly answers the query, and is explicitly supported by the referenced chunk. None - E1 Minor Omission Response is correct and source-supported, but omits minor details that do not alter the main meaning. Very Low Content E2 Partial Grounding Response is partially supported by the chunk but includes additional information not fully traceable to the source. Low Content E3 Weak Attribution Response is relevant, but the referenced chunk supports only a limited portion of the content. Moderate Content E4 Context Drift Response deviates from the focus of the query while remaining within the same general domain. Moderate Content E5 Ambiguous Response Response is overly general or ambiguous, preventing clear verification against the source. Elevated Content E6 Unsupported Claim Response contains factual claims not supported by any chunk. High Content E7 Attribution Erosion Response appears linguistically plausible but lacks valid source support in the chunks. Severe Content E8 Contradictory Content Response directly contradicts information contained in the source chunk. Severe Content E9 Abstain / No-Answer Response explicitly declines to answer or avoids providing information. Non-error (controlled) - E10 System Failure System fails to produce an evaluable output, such as empty output or processing error. Critical System Table source: Developed from the proposed evaluation framework and validated through annotation. Raw annotation data are available in Supplementary Material A. Important notes for alignment with datasets: The annotation scheme implemented in this study used the category codes E0–E10 as documented in Table 3 . In the dataset (Supplementary Material A), these codes are stored in the error_code field. It should be noted that category E7, originally labeled as "Hallucination" during the initial annotation phase, has been conceptually refined to "Attribution Erosion" based on post-hoc analysis of error patterns. This refinement captures the observation that generative errors more frequently manifest as gradual erosion of source attribution rather than complete fabrication. The raw annotation data for E7 remains unchanged and verifiable in the dataset. Similarly, E9 ("Abstain") is conceptualized as a "calibrated non-response" and a positive indicator of boundary awareness, rather than a system failure. As presented in Table 3 , this taxonomy categorizes responses across eleven distinct classes: fully supported claims, paraphrastic variations, interpretive extensions, contextual shifts, weak attributions, unsubstantiated expansions, attribution erosion, contradictions, mixed grounding failures, deliberate omissions (abstention mechanism), and system errors. The analytical separation of abstentions from attribution erosion and system errors enables evaluation of calibrated non-responses as epistemic accountability indicators rather than failure metrics. Before formal assessment, a calibration session established inter-rater consensus on interpretive thresholds without pre-determined outcome expectations. 3.6 Evaluator Selection and Structured Human Adjudication Three evaluators were selected based on their domain expertise and competencies in culturally codified meaning interpretation. All evaluators held graduate-level qualifications in cultural studies or linguistics, with specific expertise in Indonesian folklore and narrative traditions. One evaluator additionally demonstrated proficiency in natural language processing and retrieval systems. Training encompassed operationalization of the E0–E10 taxonomy, structured artifact alignment exercises, and consensus adjudication pilots. The evaluation of 61 responses yielded 183 independent annotations, with each response receiving three independent scores. In the dataset, these are recorded with unique annotation_id values linked to question_id and evaluator_id. 3.7 Selective Consensus Adjudication Following independent assessments, evaluators conducted structured human adjudication sessions. Cases were flagged for deliberation when attribution erosion typologies (E7–E10) emerged, substantial divergences among evaluators occurred, or potential grounding issues were identified. Table 4 Consensus Trigger Conditions and Resolution Structure Trigger Condition Description Action Taken Resolution Requirement Related RQ Attribution Erosion Typologies (E7–E10) Any annotation E7 to E10 Case flagged for review Unanimous agreement required RQ2 Score Divergence Severity level difference ≥ 3 Structured deliberation Artifact re-examination with full traceability RQ2 Attribution Erosion Indicators Presence of E4 classification or ambiguous cases Targeted chunk verification Evidence confirmation through source mapping RQ2 Convergent Low Severity All scores within E0–E3 range No extended deliberation Majority convergence accepted RQ2 Calibrated Non-Response E9 assigned Confirm deliberate omission Evidence absence verification RQ2 Table source: Documented consensus adjudication protocol developed for this study. Marked cases undergo collective artifact re-examination, including ground-truth documents, chunk segmentation tables, retrieved chunk identifiers, candidate responses, reference responses, and automated filtering outputs. Final decisions require unanimous agreement. In the dataset, these deliberations are documented in consensus_notes and the final decision is recorded in final_decision (accept/reject) fields. Unmarked cases are finalized based on convergence without extended deliberation. This selective mechanism operationalizes human judgment as a structured epistemic accountability framework rather than a post-hoc validation layer. This approach directly addresses RQ2 and the identified evaluative integration gaps in the Theoretical Foundation. 3.8 Traceable Artifact Preservation and Reproducible Design All evaluation components are archived and available as Supplementary Material A. To maintain anonymity during double-blind peer review, the dataset is presented with a generic title ("Supplementary Dataset for Cultural Tourism Chatbot Evaluation") and without author identifiers. The file is under embargo during review and will be made publicly available upon publication. The archive contains: 183 independent annotations (with evaluator_id, error_code, annotation_notes) 61 final consensus decisions (final_decision, consensus_notes) Curated source documents and chunk segmentation tables 61 evaluation questions with category labels (in-context/ambiguous/out-of-context) Candidate and reference responses for all questions Retrieved chunk identifiers for each response Automated filtering outputs (similarity scores, labels) Each evaluation decision is traceable to specific textual evidence and documented deliberative rationale. By preserving annotation-level and decision-level records, this framework operationalizes auditability, reproducible artifact preservation, and epistemic accountability aligned with governance standards. This approach directly addresses RQ5 and translates transparency and accountability principles into procedural implementation. 4. Results 4.1 RQ1: Measurable Levels of Cultural Faithfulness The first research question examined the discriminative capacity of tourism chatbots utilizing epistemically constrained Retrieval-Augmented Generation architectures with chunk injection as contextual anchoring mechanisms to demonstrate Cultural Faithfulness under structured granular evaluation. Of the 61 evaluated questions, 45 responses (73.8%) were accepted while 16 responses (26.2%) were rejected following documented consensus adjudication. The final decision distribution is presented in Table 5 . Table 5 Final Consensus-Based Outcomes (n = 61) Final Decision Frequency Percentage (%) Accepted 45 73.8 Rejected 16 26.2 Total 61 100 Source: Research data (can be verified in Supplementary Material A) Accepted responses included outputs categorized by grounding severity levels (E0 to E2) and calibrated non-responses (E9) in out-of-context scenarios. Rejected responses encompassed attribution erosion (E7), contradictions (E8), system failures (E10), and severe unsubstantiated elaborations. The coexistence of accepted and rejected cases indicates that the evaluation framework does not artificially inflate performance. Instead, this framework operationalizes a measurable differentiation between epistemically grounded cultural representations and invalid narrative outputs. Cultural Faithfulness is thus conceptualized not as a binary attribute but as a structured evaluative construct integrating artifact-level traceability, interpretive restraint, and attribution transparency. These findings demonstrate that in culturally codified tourism corpora, the combination of chunk-level retrieval and granular assessment enables observable and measurable narrative boundary stability preservation. 4.2 RQ2: Contribution of the Structured Human Adjudication Paradigm The second research question examines how the structured human adjudication paradigm complements artifact-level traceability systems in enhancing epistemic accountability beyond automated similarity metrics. Of the 61 evaluated cases, 16 (26.2%) were flagged for documented consensus adjudication prior to final evaluation, while 45 (73.8%) proceeded without escalation. This distribution is presented in Table 6 . Table 6 Flagged vs Non-Flagged Cases (n = 61) Category Frequency Percentage (%) Flagged for Documented Consensus Adjudication 16 26.2 Not Flagged 45 73.8 Total 61 100 Source: Research data (can be verified in Supplementary Material A) All sixteen flagged cases were rejected following artifact-level re-examination. In contrast, forty-five non-flagged cases were accepted without extended deliberation. The direct alignment between flagging status and final decisions demonstrates that structured escalation protocols effectively identified high-risk responses prior to structured human adjudication. Automated similarity metrics did not influence final outcomes. Flagged cases underwent collaborative artifact review involving curated source documents, chunk segmentation tables, retrieved chunk identifiers, and candidate-reference response comparisons. Final decisions were determined exclusively through documented consensus adjudication. These findings reveal that the structured human adjudication paradigm operates as an epistemic accountability mechanism rather than a procedural addendum. Within culturally codified narrative contexts, interpretative discrepancies may emerge despite high lexical similarity. The structured human adjudication protocols therefore function as interpretative validation frameworks, addressing subtle narrative boundary instabilities imperceptible to automated metrics. 4.3 RQ3: Distribution Patterns and Epistemic Accountability Calibration The third research question examines whether the granular scale from E0 to E10 generates meaningful distribution patterns and reflects calibrated epistemic accountability. From 183 independent annotations (61 questions × 3 evaluators), the granular score distribution is presented in Table 7 . Table 7 Distribution of Granular Scores (E0 to E10) (n = 183) Score Frequency Percentage (%) E0 57 31.1 E1 6 3.3 E2 39 21.3 E3 5 2.7 E4 27 14.8 E5 0 0.0 E6 2 1.1 E7 3 1.6 E8 3 1.6 E9 26 14.2 E10 15 8.2 Total 183 100 Source: Research data (can be verified in Supplementary Material A) Most annotations cluster within the E0 to E4 range, indicating responses predominantly anchored in source evidence with limited interpretative extensions. High-severity categories (E7 and E8) appear exclusively in rejected cases and remain infrequent, demonstrating that attribution erosion and contradictions are controlled rather than pervasive. E9 (abstention mechanism) emerges in both accepted and rejected contexts. Accepted E9 instances align with appropriate epistemic restraint for out-of-context queries, whereas rejected E9 instances reflect non-substantive non-responses or procedural insufficiencies. This distinction confirms that abstention functions as a calibrated epistemic accountability mechanism rather than a performance failure. The absence of E5 assignments reveals that ambiguous moderate attribution weaknesses are not characteristic of evaluated outputs. Annotations cluster toward clearly grounded or evidently invalid categories, indicating effective taxonomic discriminative capacity. Collectively, this distribution demonstrates layered severity differentiation and supports the claim that the E0–E10 scale enables structured epistemic accountability calibration beyond binary loyal/disloyal classifications. 4.4 RQ4: Role of Chunk-Injection in Narrative Boundary Stability The fourth research question investigates the contribution of the retrieval architecture with chunk injection to narrative consistency and verifiable grounding across the evaluated dataset. The results show that none of the accepted responses contained attribution erosion or contradictions, and all accepted responses maintained traceable alignment between generated claims and identified chunk segments within the curated corpus. This suggests that chunk-level retrieval controls inferential extension within document-bound narrative boundaries, and generative variability does not lead to catastrophic epistemic deviations in accepted outputs. To evaluate architectural consistency under identical semantic conditions, fourteen base questions were repeated, producing fourteen repeated responses (total 28 responses). The evaluator scores for each pair are presented in Table 8. For analytical clarity, severity ranges were categorized as follows: E0–E4: Low-severity grounded range E5–E6: Moderate-severity range E7–E8: Attribution erosion or contradiction range E9–E10: Abstention mechanism or system response range Table 8. Evaluator Scores, Severity Range, and Final Decisions for Repeated Question Pairs (n = 14 Pairs) Question Pair Initial Scores (E0–E10) Initial Severity Range Final Decision (Initial) Repeated Scores (E0–E10) Repeated Severity Range Final Decision (Repeated) Stability Q0–Q15 E1, E1, E2 E0–E4 Accepted E1, E2, E2 E0–E4 Accepted Stable Q02–Q16 E2, E3, E2 E0–E4 Accepted E3, E2, E3 E0–E4 Accepted Stable Q03–Q17 E4, E4, E3 E0–E4 Accepted E4, E4, E4 E0–E4 Accepted Stable Q04–Q18 E0, E0, E1 E0–E4 Accepted E0, E1, E0 E0–E4 Accepted Stable Q05–Q19 E2, E3, E2 E0–E4 Accepted E3, E4, E3 E0–E4 Accepted Stable Q06–Q20 E4, E4, E4 E0–E4 Accepted E4, E3, E4 E0–E4 Accepted Stable Q07–Q21 E9, E9, E9 E9–E10 Accepted E9, E9, E9 E9–E10 Accepted Stable Q08–Q22 E9, E9, E10 E9–E10 Accepted* E9, E9, E9 E9–E10 Accepted Stable Q09–Q23 E2, E1, E2 E0–E4 Accepted E3, E4, E3 E0–E4 Accepted Stable Q10–Q24 E3, E4, E3 E0–E4 Accepted E6, E6, E6 E5–E6 Rejected Variation Q11–Q25 E4, E4, E3 E0–E4 Accepted E6, E6, E7 E5–E8 Rejected Variation Q12–Q26 E6, E6, E6 E5–E6 Rejected E4, E4, E3 E0–E4 Accepted Variation Q13–Q27 E9, E9, E9 E9–E10 Accepted E2, E3, E2 E0–E4 Accepted Variation Q13–Q28 E6, E9, E10 E5–E10 Rejected E9, E9, E9 E9–E10 Accepted Variation Note: Q08 initial had one E10 (system failure) but was accepted after consensus as E10 was due to temporary system error; repeated generation resolved the issue. Source: Reconstruction from original research data (verifiable in Supplementary Material A) The pair Q13–Q28 (highlighted) demonstrates the stochastic nature inherent in large language models. In the initial generation (Q13), divergent evaluator scores (E6, E9, E10) indicated a complex scenario where the system generated an unsupported claim while the reference generation failed. After documented consensus adjudication, the response was rejected. In the repeated generation (Q28), both candidate and reference responses appropriately abstained, leading to acceptance. Notably, this variation did not escalate into attribution erosion (E7) or contradiction (E8) categories, showing that although generative outputs may differ, they remain constrained within epistemically controlled boundaries. Of the fourteen tested pairs: 10 pairs (71.4%) remained within the same severity range 4 pairs (28.6%) exhibited limited cross-range variation No paired instances escalated to attribution erosion (E7) or contradiction (E8) No contradictory chunk attributions were detected These results highlight the necessity of iterative testing and systematic human evaluation when assessing culturally sensitive generative systems. Consistency in this study is defined not as identical evaluator judgments, but as the preservation of severity structure and grounding alignment under repeated semantic inputs. These results indicate that chunk-level retrieval contributes to the preservation of narrative boundaries and containment of severity in culturally rooted generative tasks. 4.5 RQ5: Traceability, Auditability, and Reproducibility The fifth research question evaluates the traceability and governance alignment of the evaluation process. The study maintained comprehensive links for all 61 questions, including candidate and reference responses, chunk identifiers, filtering outputs, evaluator scores, annotation statuses, consensus documentation, and final decisions. Both rejected and accepted cases preserved annotation transparency, allowing for the reproduction of evaluation outcomes through explicit procedural links. Each evaluation decision can be traced to specific textual evidence and documented deliberative rationale via artifact-level traceability systems. The complete evaluation archive, documented in Supplementary Material A, contains: 183 independent annotations with evaluator identifiers and error codes 61 final consensus decisions with resolution metadata Curated source documents and chunk segmentation tables All candidate and reference responses with associated chunk identifiers Automated filtering outputs (similarity scores and labels) This structured documentation operationalizes governance-aligned evaluations in generative systems, which is essential in culturally sensitive tourism contexts where narrative legitimacy and epistemic accountability intersect. By preserving annotation-level and decision-level records, this framework demonstrates how abstract transparency and accountability principles can be translated into concrete, verifiable procedures. 5. Discussion This section interprets the findings presented in Section 4 in relation to the research questions and the theoretical framework established earlier. The discussion connects empirical results to prior literature, highlighting how the proposed evaluation framework advances the operationalization of Cultural Faithfulness through artifact-level traceability, severity-differentiating diagnostics, and structured human adjudication. The section is organized around the five research questions, followed by a discussion of limitations and future research directions. 5.1 RQ1: Discriminative Capacity of the Evaluation Framework The first research question examined the discriminative capacity of the proposed framework in distinguishing between grounded responses, attribution erosion, and severe grounding violations. The finding that 73.8% of responses were accepted while 26.2% were rejected confirms that the intentionally constructed dataset comprising in-context, ambiguous, and out-of-context questions successfully created challenging evaluation conditions. Rejections occurred exclusively in cases with significant grounding weaknesses, demonstrating that Cultural Faithfulness is not a binary attribute but a hierarchical construct requiring evaluation instruments sensitive to interpretive nuances. This result reinforces earlier cross-cultural NLP research showing that AI systems may produce grammatically coherent outputs that are culturally misaligned (Cao et al., 2023 ; Lee et al., 2025 ). The E0–E10 taxonomy enabled finer differentiation beyond factual accuracy, supporting the argument that cultural representation in generative systems cannot be adequately assessed through surface-level metrics alone. The dataset's success in triggering severity score variation suggests that scenario-based approaches (in-context, ambiguous, out-of-context) can serve as a model for Cultural Faithfulness evaluation in other domains. 5.2 RQ2: Role of Structured Human Adjudication The second research question investigated how structured human adjudication complements automated similarity filtering. The finding that 16 cases (26.2%) were flagged for consensus review, all of which were rejected after artifact-level re-examination, demonstrates that automated metrics alone are insufficient for detecting subtle attribution erosion and unsubstantiated expansions. Even responses with high lexical similarity scores were rejected when human evaluators identified grounding failures. These findings align with critiques of conventional human-in-the-loop models (Tschiatschek et al., 2024 ), which argue that human involvement is often superficial due to insufficient informational support. In our framework, evaluators were not merely final reviewers but decision-makers supported by rich artifacts: source chunks, traceability identifiers, similarity metrics, and evaluation records. The collective deliberation on flagged cases reflects what Salloch and Eriksen ( 2024 ) term co-reasoning , where experts jointly interpret evidence to achieve legitimate epistemic consensus. Thus, automated metrics function appropriately as an initial analytical layer that enriches human judgment rather than replacing it. This architecture operationalizes the procedural transparency demanded by AI governance principles (OECD, 2024 ), positioning structured human adjudication as an epistemic accountability mechanism rather than a post-hoc validation step. 5.3 RQ3: Diagnostic Utility of the E0–E10 Taxonomy The third research question assessed whether the E0–E10 taxonomy generates meaningful distribution patterns and reflects calibrated epistemic accountability. The concentration of scores in the E0–E4 range (70.2% of annotations) and the low frequency of E7–E8 assignments (3.2%) indicate that the taxonomy effectively captures gradations of grounding quality. The absence of E5 scores suggests a relatively strict threshold between sufficient and insufficient grounding in structured cultural narratives; deviations either remain within safe limits or fall directly into severe categories. Notably, E9 (abstention mechanism) accounted for 14.2% of annotations, appearing in both accepted and rejected contexts. Accepted E9 instances corresponded to appropriate epistemic restraint for out-of-context queries, demonstrating the system's capacity for boundary awareness (the ability to recognize knowledge limits and refrain from generating unsupported content) as conceptualized by Kapusta ( 2025 ). Rejected E9 instances, by contrast, reflected procedural insufficiencies or non-substantive non-responses. This distinction confirms that abstention functions as a calibrated epistemic accountability indicator rather than a performance failure. The taxonomy thus serves dual purposes: as a diagnostic tool for identifying error types and as a calibration instrument that distinguishes between generative inability and intentional restraint. By enabling evaluators to verify evidence absence through artifact-level traceability, the framework renders abstention decisions accountable and transparent. 5.4 RQ4: Contribution of Chunk Injection to Narrative Boundary Stability The fourth research question explored the role of chunk injection in maintaining narrative boundary stability during iterative generation. The repeated testing of 14 question pairs revealed that 71.4% remained within the same severity range, with no instances escalating to attribution erosion (E7) or contradiction (E8). This stability indicates that explicit chunk injection into the model context acts as an epistemic constraint, limiting generative space and preventing extrapolation beyond retrieved evidence. These findings support Hickey's (2026) concept of constitutive knowledge sources : epistemic trust is fostered not through algorithmic transparency but through authoritative, verifiable knowledge sources. Injected chunks serve precisely this function, transforming retrieval from probabilistic similarity matching to deterministic evidence provision. The observed stability also aligns with Kapusta's (2025) vision of two-level transparency, where high-level reasoning patterns and detailed factor explanations remain traceable. Annotators could consistently verify outputs across generations because chunk-level traceability provided sufficient evidence for in-depth validation. The Q13–Q28 pair illustrates the stochastic nature of large language models while also demonstrating the effectiveness of epistemic constraints. Although the initial generation produced divergent scores (E6, E9, E10) leading to rejection, the repeated generation appropriately abstained (E9) and was accepted. Crucially, this variation did not escalate into attribution erosion or contradiction, confirming that even when outputs differ, they remain bounded by the retrieved evidence. 5.5 RQ5: Traceability, Auditability, and Governance Alignment The fifth research question evaluated how artifact preservation operationalizes transparency and accountability. The complete evaluation archive (Supplementary Material A) enables full reconstruction of every decision, from source chunks to consensus metadata. This exemplifies the OECD AI Principles' (2024) traceability-as-accountability instrument, allowing outcome analysis and inquiry resolution. In culturally sensitive domains where misrepresentation can harm heritage-owning communities, auditability becomes essential. Our framework operationalizes Shneiderman's (2022) Human-Centered AI vision by embedding human control and transparent reporting into the evaluation design. Annotators function not merely as validators but as holders of epistemic authority whose decisions are documented and accountable. This aligns with principles of moral authorship (Cristofaro & Bañón Gomis, 2026 ), where final decisions remain under human purview to preserve autonomy and responsibility. By providing an open dataset and comprehensive documentation, this research contributes to open science initiatives and enables replication, verification, and extension by the research community. The framework demonstrates how abstract governance principles (transparency, accountability, and reproducibility) can be translated into concrete, verifiable evaluation protocols. 5.6 Limitations and Future Research Several limitations should be acknowledged. First, the study employed a single cultural corpus (the Ande-Ande Lumut narrative). While this enabled controlled experimentation, generalization to other cultural contexts requires cross-corpus validation with diverse folklore, traditions, and narrative structures. Future research should apply the framework to multiple cultural datasets to assess its adaptability and identify culture-specific error patterns. Second, the evaluator panel consisted of three domain experts. Although inter-rater consensus was achieved through calibration and documented consensus adjudication, larger panels would enhance statistical robustness and allow for more nuanced analysis of evaluator effects. Future studies could incorporate evaluators from different cultural backgrounds to examine how interpretive frameworks influence attribution erosion judgments. Third, while abstention (E9) is conceptualized as a positive indicator of boundary awareness, practical deployment may require balancing epistemic caution with user expectations. Users interacting with tourism chatbots typically anticipate informative responses; excessive abstention could degrade user experience. Future research should explore optimal trade-offs between epistemic restraint and user satisfaction, potentially developing adaptive abstention policies. Fourth, the current taxonomy, while granular, may benefit from refinement for partial automation. Developing concise, computationally detectable proxies for certain error categories could streamline evaluation without sacrificing diagnostic precision. Additionally, integrating alternative traceability methods such as GraphRAG (Edge et al., 2024) could address complex multi-hop questions requiring reasoning across distributed knowledge sources. Finally, longitudinal studies are needed to assess how narrative boundary stability evolves over multiple interaction turns and whether repeated exposure to similar queries leads to attribution erosion accumulation. Such research would inform the design of self-correcting generative systems capable of maintaining cultural faithfulness over extended deployments. 5.7 Summary Collectively, these findings demonstrate that Cultural Faithfulness can be systematically operationalized through the integration of artifact-level traceability, severity-differentiating diagnostics, structured human adjudication, and reproducible artifact preservation. The proposed framework advances evaluation practices beyond surface-level similarity metrics toward a unified model of epistemic accountability for generative systems in tourism and other culturally sensitive domains. By bridging smart tourism research, retrieval-augmented language modeling, structured human evaluation, and operational AI governance, this study provides a replicable template for assessing and ensuring cultural faithfulness in AI-mediated communication. 6. Conclusion This study introduces and empirically validates a traceable structured human adjudication framework for assessing Cultural Faithfulness in epistemically constrained Retrieval-Augmented Generation systems for tourism chatbots. Building upon UNESCO's (2003) normative framework recognizing oral traditions as intangible cultural heritage and OECD's (2024) governance principles positioning artifact-level traceability as an accountability instrument, this research addresses an urgent need for evaluation methodologies beyond conventional automated metrics. Empirical findings from 61 structured questions and 183 independent artifact-level traceability annotations reveal that 73.8% of responses were accepted after documented consensus adjudication, while 26.2% were rejected due to attribution erosion or severe grounding violations. Crucially, repeated testing showed no escalation into attribution erosion (E7) or contradiction (E8) categories, indicating the success of the boundary awareness mechanism embedded in the framework through chunk injection as contextual anchoring mechanisms. Fundamental Contributions This study presents four fundamental contributions that collectively shift the paradigm of AI evaluation from a technical-binary approach to epistemic accountability. First, the study articulates the role of automated metrics as an analytical signal rather than a final determinant. Unlike conventional practices that treat ROUGE, BERTScore, or cosine similarity as proxies for truth, our framework positions these metrics as generators of initial diagnostic artifacts that enrich information for structured human adjudication. This design acknowledges the limitations of automated metrics in capturing nuanced fidelity aspects (Maynez et al., 2020 ) while empowering practical decision-makers with comprehensive information to fulfill their ethical and epistemic roles (Tschiatschek et al., 2024 ). Automated similarity filtering thus functions as a complementary layer within artifact-level traceability systems, not as a grounding validator. Second, the study redefines abstention and rejection as indicators of epistemic accountability. Within our framework, decisions to reject responses or invoke the abstention mechanism are not interpreted as system failures but as positive evidence of boundary awareness (the ability to recognize knowledge limits). This aligns with the vision of symbiotic epistemology (Kapusta, 2025 ), where reliable epistemic agents are those capable of acknowledging their ignorance, for instance by responding with " Maaf, saya tidak tahu " (the local expression for "I do not know"), rather than providing answers with false confidence. In culturally codified contexts, abstention becomes an ethical act that prevents potential harm from cultural misrepresentation (Floridi et al., 2018 ). The deliberate omission mechanism thus operationalizes epistemic accountability as a measurable construct. Third, the study operationalizes reproducible artifact preservation across the entire pipeline as a foundation for accountability. The framework systematically produces and preserves artifacts at every stage: from injected source chunks, through severity-differentiating diagnostics (E0 to E10), attribution erosion indicators, to final documented consensus adjudication decisions. This artifact collection enables artifact-level traceability, auditability, and epistemic accountability of every decision. It represents a concrete implementation of transparency as an accountability instrument (OECD, 2024 ) and constitutive knowledge as the foundation for epistemic trust in opaque systems (Hickey, 2026 ). The reproducible artifact preservation framework ensures that every evaluation decision is traceable to specific textual evidence through artifact-level traceability systems. Fourth, the study introduces a granular error taxonomy from E0 to E10 as a collaborative reasoning instrument for structured human adjudication. This severity-differentiating taxonomy elevates evaluation from binary judgments of correctness to rich epistemic diagnostics of narrative boundary stability. It enables annotators to function as co-reasoners (Salloch & Eriksen, 2024 ) who collaboratively interpret evidence, categorize attribution erosion typologies, and achieve legitimate documented consensus. This process generates moral authorship (Cristofaro & Bañón Gomis, 2026 ), where final decisions remain genuinely human, not merely formalities within algorithmic loops. The taxonomy thus serves as both diagnostic instrument and governance mechanism for culturally rooted narrative agents. Implications and Future Research This study not only proposes an alternative evaluation method but establishes a novel epistemic accountability paradigm for generative AI. The paradigm emphasizes that in culturally sensitive domains like tourism, Cultural Faithfulness cannot be reduced to similarity scores or statistical thresholds. It necessitates active human engagement as epistemic authorities, supported by artifact-level traceability systems and structured collaborative reasoning processes. The finding that 26.2% of responses were rejected demonstrates that Cultural Faithfulness constitutes a dimension far more demanding than mere factual accuracy. This figure does not signify system weakness but rather reflects the complexity of culturally codified heritage resisting simplification into vector representations. Crucially, it reminds us that AI development for cultural tourism must collaborate with cultural custodians, not merely rely on technical optimization. The rejection rate operationalizes the gap between technical performance and epistemic accountability. Future research directions include: Taxonomy transferability across diverse cultural domains to validate the severity-differentiating diagnostics framework. Optimizing the balance between automated similarity filtering and structured human adjudication to enhance efficiency without compromising epistemic accountability. Integrating with alternative traceability architectures such as GraphRAG (Edge et al., 2024) to address complex multi-hop queries while maintaining narrative boundary stability. Developing governance frameworks that embed reproducible artifact preservation into institutional practices for cultural heritage protection. Most importantly, this framework offers a governance model where technology does not replace but rather strengthens human roles as cultural heritage stewards in the digital era. By operationalizing transparency and accountability principles through artifact-level traceability systems, structured human adjudication paradigms, and reproducible artifact preservation mechanisms, this research contributes to a broader vision—developing generative systems that are not only technically intelligent but also culturally faithful and epistemically accountable. Declarations Author Contribution All authors contributed to the study conception and design. Hindun Nurhidayati was responsible for the cultural tourism conceptual framework and dataset preparation, drawing upon the foundational book Dilema penutur dalam pariwisata budaya co-authored by all contributors. The communication and language aspects were overseen by Zulin Nurchayati. Sugiarto, as an independent researcher, led the technical implementation of AI models, retrieval architecture design, and data management. All authors read and approved the final manuscript. Acknowledgement The authors express their gratitude to colleagues for their valuable discussions and support throughout this research. Special appreciation is also extended to the evaluators who participated in the annotation and adjudication processes. Data Availability The dataset generated and analyzed during the current study is owned by the authors and has been deposited in the Zenodo repository under a temporary embargo to ensure double blind review. The dataset will be made publicly available upon publication at: https://doi.org/10.5281/zenodo.18603955During the peer review process, the dataset is provided as Supplementary Material A (file name: Supplementary_Material_A_Cultural_Faithfulness_Dataset.pdf) under a generic title that does not reveal author identities. This material contains all evaluation artifacts necessary for replicating the study.Note to the EditorThe authors wish to inform the editor that the two references cited in the manuscript with the author names Nurhidayati, Nurchayati, & Sugiarto (2025, 2026) are works authored by the same research team. These publications, comprising a book and a conceptual preprint, represent a deliberate research trajectory that establishes the theoretical and empirical foundation for the present study. This disclosure is made to avoid any perception of self‑plagiarism and to ensure full transparency. Funding Declarations The authors declare that no funds, grants, or other support were received during the preparation of this manuscript. References Abid A, Farooqi M, Zou J (2021) Persistent anti-Muslim bias in large language models. In: Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society , pp. 298–306. https://doi.org/10.1145/3461702.3462624 Amershi S, Weld D, Vorvoreanu M, Fourney A, Nushi B, Collisson P, Suh J, Iqbal S, Bennett PN, Inkpen K, Teevan J, Kikin-Gil R, Horvitz E (2019) Guidelines for human-AI interaction. In: Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems , Paper No. 3, pp. 1–13. https://doi.org/10.1145/3290605.3300233 Artstein R, Poesio M (2008) Inter-coder agreement for computational linguistics. Comput Linguistics 4555–596. https://doi.org/10.1162/coli.07-034-R2 Asai A, Salehi M, Peters ME, Hajishirzi H (2022) ATTEMPT: Parameter-efficient multi-task tuning via attentional mixtures of soft prompts. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , 6655–6672. https://doi.org/10.18653/v1/2022.emnlp-main.446 Bender EM, Gebru T, McMillan-Major A, Shmitchell S (2021) On the dangers of stochastic parrots: Can language models be too big? Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , 610–623. https://doi.org/10.1145/3442188.3445922 Birhane A, Kalluri P, Card D, Agnew W, Dotan R, Bao M (2022) The values encoded in machine learning research. In: Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (FAccT '22) , pp. 173–184. https://doi.org/10.1145/3531146.3533083 Blodgett SL, Barocas S, Daumé H III, Wallach H (2020) Language (technology) is power: A critical survey of bias in NLP. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 5454–5476. https://doi.org/10.18653/v1/2020.acl-main.485 Borgeaud S, Mensch A, Hoffmann J, Cai T, Rutherford E, Casas K, Sifre L (2022) Improving language models by retrieving from trillions of tokens. Proceedings of the 39th International Conference on Machine Learning , PMLR 162, 2206–2240. https://proceedings.mlr.press Campbell JL, Quincy C, Osserman J, Pedersen OK (2013) Coding in-depth semistructured interviews: Problems of unitization and intercoder reliability and agreement. Sociol Methods Res 3294–320. https://doi.org/10.1177/0049124113500475 Cao Y, Zhou L, Lee S, Cabello L, Chen M, Hershcovich D (2023) Assessing cross-cultural alignment between ChatGPT and human societies: An empirical study. *Proceedings of the First Workshop on Cross-Cultural Considerations in NLP (C3NLP)*. https://aclanthology.org/2023.c3nlp-1.3 Caric H, Mandić A, Sever I (2026) A six-phase AI-expert framework for evaluating policy coherence in sustainable tourism. J Inform Technol Tourism 6. https://doi.org/10.1007/s40558-025-00352-0 . *28* Cohen J (1960) A coefficient of agreement for nominal scales. Educ Psychol Meas 20(1):37–46. https://doi.org/10.1177/001316446002000104 Cristofaro M, Bañón Gomis G (2026) Dancing with the algorithm: A framework to navigate knowledge and autonomy in AI-assisted managerial decisions. J Knowl Manage. https://doi.org/10.1108/JKM-06-2025-0870 Cruz M, Jardim B, de Castro Neto M (2025) Lisa: A touristic chatbot for Lisbon. J Inform Technol Tourism 1153–1183. https://doi.org/10.1007/s40558-025-00339-x . *27* Edge D, Trinh H, Cheng N, Bradley J, Chao A, Mody A, Truitt S, Metropolitansky D, Ness RO, Larson J (2025) From local to global: A graph RAG approach to query-focused summarization (Version 2). arXiv. https://doi.org/10.48550/arXiv.2404.16130 Fleiss JL (1971) Measuring nominal scale agreement among many raters. Psychological Bulletin , *76*(5), 378–382. https://doi.org/10.1037/h0031619 Floridi L, Cowls J, Beltrametti M et al (2018) AI4People—An ethical framework for a good AI society: Opportunities, risks, principles, and recommendations. Minds and Machines , *28*, 689–707. https://doi.org/10.1007/s11023-018-9482-5 Gao Y, Xiong Y, Gao X, Jia H, Pan Y, Bi Y, Dai Y, Sun J, Wang H, Wang H (2024) Retrieval-augmented generation for large language models: A survey. arXiv preprint. https://arxiv.org/abs/2312.10997 Gretzel U (2011) Intelligent systems in tourism: A social science perspective. Annals Tourism Res 38(3):757–781. https://doi.org/10.1016/j.annals.2011.04.014 Gretzel U, Sigala M, Xiang Z, Koo C (2015) Smart tourism: Foundations and developments. Electron Markets 25(3):179–188. https://doi.org/10.1007/s12525-015-0196-8 Guu K, Lee K, Tung Z, Pasupat P, Chang MW (2020) Retrieval augmented language model training. In Proceedings of the 37th International Conference on Machine Learning (Vol. 119, pp. 3929–3938). PMLR. https://proceedings.mlr.press Hershcovich D, Frank S, Lent H, de Lhoneux M, Abdou M, Brandl S, Bugliarello E, Piqueras C, Chalkidis L, Cui I, Fierro R, Margatina C, Rust K, P., Søgaard A (2022) Challenges and strategies in cross-cultural NLP. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics , 6997–7013. https://doi.org/10.18653/v1/2022.acl-long.482 Hickey C (2026) Constitutive knowledge sources: An institutional approach to epistemic trust in opaque AI systems. AI and Ethics , *6*, 58. https://doi.org/10.1007/s43681-025-00930-2 Hofstede G, Hofstede GJ, Minkov M (2010) Cultures and organizations: Software of the mind: Intercultural cooperation and its importance for survival, 3rd edn. McGraw-Hill Ji Z, Han S, Yu S (2023) Survey of hallucination in natural language generation. ACM-CSUR 55(12):1–38. https://doi.org/10.1145/3571730 Jobin A, Ienca M, Vayena E (2019) The global landscape of AI ethics guidelines. Nat Mach Intell 9389–399. https://doi.org/10.1038/s42256-019-0088-2 Kapusta J (2025) SynLang and symbiotic epistemology: A manifesto for conscious human-AI collaboration. arXiv preprint. https://arxiv.org/abs/2507.21067 Kim JH, Kim J, Park J, Kim C, Jhang J, King B (2025) When ChatGPT gives incorrect answers: The impact of inaccurate information by generative AI on tourism decision-making. J Travel Res 64(1):51–73. https://doi.org/10.1177/00472875231212996 Krippendorff K (2019) Content analysis: An introduction to its methodology (4th ed.). SAGE Publications. https://doi.org/10.4135/9781071878781 Landis JR, Koch GG (1977) The measurement of observer agreement for categorical data. Biometrics *33* 1159–174. https://doi.org/10.2307/2529310 Laux J, Wachter S, Mittelstadt B (2024) Trustworthy artificial intelligence and the European Union AI act: On the conflation of trustworthiness and acceptability of risk. Regul Gov 18(1):3–32. https://doi.org/10.1111/rego.12512 Lee N, Bang Y, Madotto A, Fung P (2025) Cultural alignment test (Hofstede's CAT): Revealing the struggle of Large Language Models in comprehending cultural values. On Proceedings of the 31st International Conference on Computational Linguistics (COLING 2025) . Association for Computational Linguistics. https://aclanthology.org/2025.coling-main.567/ Lewis P, Perez E, Piktus A, Petroni F, Karpukhin V, Goyal N, Küttler H, Lewis M, Yih WT, Rocktäschel T, Riedel S, Kiela D (2020) Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474. proceedings.neurips.cc Lin C-Y (2004) ROUGE: A package for automatic evaluation of summaries. Proceedings of the Association for Computational Linguistics Workshop , 74–81. https://aclanthology.org/W04-1013/ Liu CC, Gurevych I, Korhonen A (2025) Culturally aware and adapted NLP: A taxonomy and a survey of the state of the art. Trans Association Comput Linguistics 13:652–689. https://doi.org/10.1162/tacl_a_00760 Lucy L, Bamman D (2021) Gender and representation bias in GPT-3 generated stories. In: Proceedings of the Third Workshop on Narrative Understanding , pp. 48–55. Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.nuse-1.5 Maynez J, Narayan S, Bohnet B, McDonald R (2020) On faithfulness and factuality in abstractive summarization. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 1906–1919. https://aclanthology.org/2020.acl-main.173 Mittelstadt BD, Allo P, Taddeo M et al (2016) The ethics of algorithms. Big Data Soc. *3*(2 https://doi.org/10.1177/2053951716679679 Mohamed S, Png MT, Isaac W (2020) Decolonial AI: Decolonial theory as sociotechnical foresight in artificial intelligence. Philos Technol 4659–684. https://doi.org/10.1007/s13347-020-00405-8 Nurhidayati H, Nurchayati Z, Sugiarto (2025) Dilema penutur dalam pariwisata budaya. Deepublish Nurhidayati H, Nurchayati Z, Sugiarto (2026) Lima teori konseptual untuk integrasi AI dan komunikasi budaya dalam pariwisata. Preprint at Zenodo . https://doi.org/10.5281/zenodo.18335863 OECD (2019) OECD principles on artificial intelligence. OECD Publishing. https://oecd.ai/en/ai-principles OECD (2024) OECD Artificial Intelligence Principles 2024 . OECD Publishing. https://oecd.ai/en/ai-principles Ren J, Xu Y, Wang X, Li W, Wang A, Ma W, Liu Y (2025) Towards transparent RAG: Fostering evidence traceability in LLM generation via reinforcement learning. arXiv preprint. https://doi.org/10.48550/arXiv.2505.13258 Salemi A, Zamani H (2024) Evaluating retrieval quality in retrieval-augmented generation. Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval , 2395–2400. https://doi.org/10.1145/3626772.3657957 Salloch S, Eriksen A (2024) What are humans doing in the loop? Co-reasoning and practical judgment when using machine learning-driven decision aids. Am J Bioeth 24(9):67–78. https://doi.org/10.1080/15265161.2024.2353800 Shneiderman B (2022) Human-centered AI. Oxford University Press. https://doi.org/10.1093/oso/9780192845290.001.0001 Sigala M (2018) New technologies in tourism: From multi-disciplinary to anti-disciplinary advances and trajectories. Tourism Management Perspectives, *25*, 151–155. https://doi.org/10.1016/j.tmp.2017.12.003 Stilgoe J, Owen R, Macnaghten P (2013) Developing a framework for responsible innovation. Research Policy , *42*(9), 1568–1580. https://doi.org/10.1016/j.respol.2013.05.008 Tafjord O, Dalvi B, Clark P (2021) ProofWriter: Generating implications, proofs, and abductive statements over natural language. Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , 3621–3634. https://doi.org/10.18653/v1/2021.findings-acl.317 Ting-Toomey S (2015) Intercultural and intergroup communication competence: Toward an integrative perspective. In G. Rickheit & H. Strohner (Eds.), Handbook of communication competence (pp. 503–538). De Gruyter Mouton. https://doi.org/10.1515/9783110317459-021 Tsamados A, Aggarwal N, Cowls J, Morley J, Roberts H, Taddeo M, Floridi L (2022) The ethics of algorithms: key problems and solutions. AI & Society, *37*, pp 215–230. 1 https://doi.org/10.1007/s00146-021-01154-8 Tschiatschek S, Stamboliev E, Schmude T, Coeckelbergh M, Koesten L (2024) Challenging the human-in-the-loop in algorithmic decision-making. arXiv preprint. https://doi.org/10.48550/arXiv.2405.10706 UNESCO (2003) Convention for the safeguarding of the intangible cultural heritage. UNESCO. https://ich.unesco.org/en/convention UNESCO (2025) AI and the future of education: Disruptions, dilemmas and directions. UNESCO Publishing. https://unesdoc.unesco.org/ark:/48223/pf0000395236 Zhang T, Kishore V, Wu F, Weinberger KQ, Artzi Y (2020) BERTScore: Evaluating text generation with BERT. In Proceedings of the 8th International Conference on Learning Representations (ICLR 2020). OpenReview.net. https://openreview.net/forum?id=SkeHuCVFDr Additional Declarations No competing interests reported. Supplementary Files SupplementaryMaterialAITT.pdf Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8994226","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":607686787,"identity":"587ac3f8-7428-44cc-9efe-c016bc118133","order_by":0,"name":"Hindun Nurhidayati -","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABRklEQVRIie2QQWvCMBSAUwLp5dVzRMf+QkCYDIdl/6ShUE+KYyADCwaEepHt6mDoX5gXzxuB9CKeB7tUBE87DHYRxmBpR7FuuvNg/SDk5b33keQhlJPzNyHIIfFuCIyu4mOaN792QA8HFBIr86yCf1GSNr2wEWQLOA12lSptrqOoU+c3k77AF2PZKtDm41vbr9sFjCO08VG5sKucjhpV5ixcPlL6YbczeUloyy2NlMsDTJgxVAiKuwp78gjVVS5ixZpJHsCclSyBHf0XhiyBgO1VenySKHeJUnm3RM8m2Hw1Pg4qkt8nitCKOTzRt0gjwMDwvlvma0KdRViZKi4kqIZWrr0aqFC/FtqyrOiPv4QeKW463aOxDJcr8Gt80rfkM/hd+3gwmC5f/DP728RSzsV2/rBN6wzd26+xMzEcasrJycn5p3wC8PJu0mS1NwgAAAAASUVORK5CYII=","orcid":"","institution":"Pancasila University","correspondingAuthor":true,"prefix":"","firstName":"Hindun","middleName":"Nurhidayati","lastName":"-","suffix":""},{"id":607686788,"identity":"093f0c62-5fe0-4268-b9df-9cc5b5477cd3","order_by":1,"name":"Zulin Nurchayati -","email":"","orcid":"","institution":"Universitas Merdeka Madiun","correspondingAuthor":false,"prefix":"","firstName":"Zulin","middleName":"Nurchayati","lastName":"-","suffix":""},{"id":607686789,"identity":"e3132ff4-a409-4364-a0e4-a18f5410f1de","order_by":2,"name":"Sugiarto -","email":"","orcid":"","institution":"","correspondingAuthor":false,"prefix":"","firstName":"Sugiarto","middleName":"","lastName":"-","suffix":""}],"badges":[],"createdAt":"2026-02-28 10:08:18","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8994226/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8994226/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":104887677,"identity":"b7c3bb8e-981f-4cc8-8535-d3eaffefb778","added_by":"auto","created_at":"2026-03-18 10:12:31","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":4405192,"visible":true,"origin":"","legend":"\u003cp\u003eStructured Human Adjudication Architecture for Cultural Faithfulness, Showing Retrieval-Augmented Generation (RAG) Components\u003c/p\u003e","description":"","filename":"Figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-8994226/v1/835c15af0589bac3788cb9e6.png"},{"id":106912431,"identity":"c1d532c8-69cc-419d-a5c9-ced0abc1c4a5","added_by":"auto","created_at":"2026-04-14 16:56:22","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":5633236,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8994226/v1/549d7aae-69ee-43b9-98dd-1fbec8c4352a.pdf"},{"id":104887710,"identity":"41808ee0-7475-4884-9138-8b6eea5ad660","added_by":"auto","created_at":"2026-03-18 10:12:33","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":8329733,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryMaterialAITT.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8994226/v1/94bc7adc8de0e999d76c1574.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Cultural Faithfulness in Tourism Chatbots: A Structured Human Adjudication Framework for Traceable Retrieval-Augmented Generation","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eThe digital transformation in tourism has driven the adoption of large language model (LLM) chatbots, which now serve dual roles as information providers and narrative agents. These systems influence how tourists engage with folklore, local history, and traditional practices, raising critical questions about their cultural responsibilities. When a chatbot represents a culture, factual accuracy is insufficient; cultural communication requires fidelity to source narratives, contextual sensitivity, and epistemic accountability for knowledge representation.\u003c/p\u003e \u003cp\u003eInternational frameworks emphasize the protection of intangible cultural heritage. The 2003 UNESCO Convention identifies oral traditions as vital domains requiring preservation through intergenerational transmission and protection from distortion (UNESCO, \u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e2003\u003c/span\u003e). In artificial intelligence contexts, this raises a contemporary challenge: how to ensure digital agents preserve cultural values when communicating folklore? UNESCO (\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e2025\u003c/span\u003e) warns that algorithms cannot fully replace human values, underscoring the need for cultural guardians to retain control over heritage representation.\u003c/p\u003e \u003cp\u003eRetrieval-Augmented Generation (RAG) architectures aim to improve generative models' factual precision by conditioning outputs on external documents (Lewis et al., \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Gao et al., \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). However, evaluation remains dominated by automated metrics like ROUGE (Lin, \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2004\u003c/span\u003e) and BERTScore (Zhang et al., \u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e2020\u003c/span\u003e), which assess surface-level similarity rather than epistemic alignment with source evidence. Cross-cultural NLP research reveals that large language models may produce grammatically coherent outputs misaligned with local community values, highlighting cultural gaps in current AI systems (Cao et al., \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2023\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThe limitations inherent in digital storytelling systems carry significant implications for cultural tourism. Subtle interpretive expansions\u0026mdash;such as the inclusion of unverified details, oversimplification of narrative arcs, or moral reframing\u0026mdash;pose a risk of displacing narrative authority from cultural communities to algorithmic systems. This phenomenon becomes particularly critical when chatbots function as digital narrators within culturally rich environments, where insufficient accountability mechanisms could fundamentally alter knowledge transmission dynamics (Nurhidayati et al., \u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e2025\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThe central issue transcends basic attribution erosion detection, revealing a systemic lack of diagnostic frameworks for multi-layered grounding failures. These failures manifest as attribution erosion, unsubstantiated expansions, or outright factual contradictions. While prior research has addressed hallucination detection (Ji et al., \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2023\u003c/span\u003e) and the impact of inaccurate generative outputs on consumer trust in tourism (Kim et al., \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e2025\u003c/span\u003e), three critical structural deficiencies persist:\u003c/p\u003e \u003cp\u003eFirst, current evaluation protocols rarely combine artifact-level traceability with systematic human oversight. Automated evaluation tools may identify semantic inconsistencies but remain ineffective at detecting attribution errors, unsubstantiated expansions, or boundary shifts in narrative frameworks.\u003c/p\u003e \u003cp\u003eSecond, Human-in-the-Loop (HITL) implementations predominantly focus on model optimization rather than epistemic validation. In conventional systems, human review operates as an after-the-fact verification process rather than a structured evaluation phase with defined diagnostic criteria. Contemporary critiques highlight that mere human involvement does not ensure ethical or accurate system behavior; this requires explicit role delineation and comprehensive information support for adjudicators (Tschiatschek et al., \u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e2024\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThird, governance principles like transparency and accountability are frequently stated in abstract terms without concrete implementation through artifact-level traceability or documented consensus adjudication protocols. The 2024 OECD AI Principles explicitly recognize traceability as an accountability mechanism for outcome analysis and inquiry resolution (OECD, \u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). Despite this, human-mediated adjudication\u0026mdash;intended to operationalize these principles\u0026mdash;remains underdeveloped in Retrieval-Augmented Generation (RAG) systems.\u003c/p\u003e \u003cp\u003eThe proposed evaluation framework for Retrieval-Augmented Generation (RAG) in tourism chatbots addresses critical gaps by integrating four methodological innovations. First, source traceability is preserved through artifact-level chunk injection, where cultural narratives are explicitly embedded into the model context. This transforms retrieval from probabilistic similarity matching to deterministic evidence provision, a foundational requirement for epistemic accountability (Hickey, \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2026\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eSecond, a granular error classification system (E0\u0026ndash;E10) was developed to assess grounding severity. The taxonomy differentiates between fully supported responses (E0), minor omissions (E1), partial grounding with untraceable additions (E2), and severe misrepresentations like contradictions (E8) or attribution erosion (E7). This replaces binary true/false evaluations with a diagnostic framework that captures nuanced error types.\u003c/p\u003e \u003cp\u003eThird, automated metrics such as BERTScore and ROUGE-L function as analytical supports rather than final judgments. These generate diagnostic artifacts that inform human evaluators, addressing critiques about the misalignment between automated metrics and human assessments of faithfulness (Maynez et al., \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e2020\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eFinally, epistemic validation is formalized through documented consensus adjudication. Annotators collectively decide on acceptance, rejection, or abstention using co-reasoning (Salloch \u0026amp; Eriksen, \u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). The abstention mechanism explicitly signals boundary awareness, aligning with symbiotic epistemology (Kapusta, \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e2025\u003c/span\u003e) and ethical obligations to avoid cultural misrepresentation (Floridi et al., \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2018\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThe proposed evaluation framework transcends conventional output scoring by functioning as a systematic methodology to assess whether retrieval mechanisms impose epistemic constraints on generative processes. Cultural Faithfulness is implemented through three core components: components: artifact-level traceability for source verification, severity-differentiating diagnostics for error classification, and reproducible artifact preservation for accountability. To ensure methodological rigor, the evaluation dataset has been archived as an open research artifact (see Supplementary Material A), containing structured queries, granular annotation logs, chunk-level mapping data, and consensus metadata.\u003c/p\u003e \u003cp\u003eThis study investigates five central research questions (RQ):\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eWhat is the discriminative capacity of the evaluation framework in distinguishing grounded responses from attribution erosion and severe grounding violations?\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eHow does automated similarity filtering complement structured human adjudication rather than supplant it?\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eIn what ways does the E0\u0026ndash;E10 taxonomy function as a diagnostic tool for assessing grounding quality?\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eWhat role does chunk-injection retrieval play in maintaining narrative boundary stability during iterative generation processes?\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eHow does artifact preservation operationalize transparency and accountability in culturally situated chatbot systems?\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cp\u003eThis study presents four key contributions to the fields of smart tourism and responsible artificial intelligence. First, it expands the understanding of Cultural Faithfulness as an evaluative framework specifically tailored for tourism chatbots. By redefining chatbots as culturally embedded narrative facilitators rather than mere information providers, the research connects digital tourism scholarship with cross-cultural natural language processing (NLP) assessment methodologies.\u003c/p\u003e \u003cp\u003eSecond, the study proposes a spectrum-based diagnostic classification system (E0\u0026ndash;E10) to evaluate attribution erosion in retrieval-augmented generation (RAG) systems. This taxonomy moves beyond simplistic true/false accuracy metrics by systematically categorizing response types including grounded assertions, contextual transitions, alignment errors, attribution erosion, contradictory statements, deliberate omissions, and systemic malfunctions.\u003c/p\u003e \u003cp\u003eThird, the research implements narrative boundary stability through a structured human adjudication architecture incorporating chunk injection. This approach ensures artifact-level traceability by directly linking generated content to source text segments while maintaining evaluation records. The methodology offers a replicable template for assessing culturally sensitive generative systems.\u003c/p\u003e \u003cp\u003eFourth, the framework advances tourism AI governance by converting abstract transparency and accountability principles into actionable evaluation protocols. Through version-controlled dataset management, documented consensus adjudication practices, and annotation-level record preservation, the study demonstrates practical governance implementation in tourism AI applications. These contributions collectively forge an interdisciplinary nexus between smart tourism research, retrieval-augmented language modeling, structured human evaluation, and operational AI governance.\u003c/p\u003e"},{"header":"2. Literature Review and Theoretical Background","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1 Digital Tourism, Cultural Mediation, and Epistemic Accountability\u003c/h2\u003e \u003cp\u003eDigital transformation has permeated nearly all sectors, with tourism communication undergoing significant reconfiguration. Smart systems now mediate interactions between destinations and travelers, transforming tourism platforms from passive information repositories into culturally rooted narrative agents. Gretzel et al. (\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2015\u003c/span\u003e) documented how this evolution reconfigures communication paradigms. Smart tourism ecosystems have institutionalized data-driven interactivity, algorithmic personalization, and the digital co-creation of experiences (Sigala, \u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e2018\u003c/span\u003e). Recent studies emphasize the integration of AI into destination management practices, including chatbots that operationalize transparency and accountability (Caric et al., \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e2026\u003c/span\u003e) and culturally embedded facilitators within smart tourism infrastructures (Cruz et al., \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e2025\u003c/span\u003e). These developments demonstrate that algorithmic mediation has transitioned from peripheral to structurally embedded within tourism communication systems.\u003c/p\u003e \u003cp\u003eDespite these advancements, smart tourism research predominantly emphasizes competitiveness, operational efficiency, personalization, and sustainability. Consequently, systematic attention to what can be termed \"epistemic accountability\" \u0026mdash;the responsibility of a knowledge system to accurately represent its sources and the boundaries of its own knowledge \u0026mdash;within cultural representation remains relatively limited. This gap is particularly consequential given the normative obligations attached to the protection of intangible cultural heritage, which demand authenticity, contextual coherence, and faithful community representation (UNESCO, \u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e2003\u003c/span\u003e). When generative systems, such as tourism chatbots, construct cultural narratives, they inevitably assume an interpretive authority that extends beyond mere factual transmission. They make choices about what to include, exclude, and emphasize, thereby shaping the user's understanding of the culture.\u003c/p\u003e \u003cp\u003eIn the domain of digital tourism, intelligent systems function as critical mediators that fundamentally alter how travelers perceive and interact with cultural heritage (Gretzel, \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e2011\u003c/span\u003e). From the perspective of human-centered AI and responsible innovation, such systems must be evaluated not only by their performance metrics but also by their alignment with broader social and ethical expectations (Shneiderman, 2020; Floridi et al., \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2018\u003c/span\u003e). Beyond mere technical performance, the development of culturally faithful chatbots must be grounded in the Responsible Innovation (RI) framework, which advocates for a proactive and reflective approach to technological advancement (Stilgoe et al., \u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e2013\u003c/span\u003e). This evaluation must scrutinize not just whether the information is factually correct, but whether the system's mode of presentation and its handling of uncertainty respect the epistemic integrity of the source culture. Consequently, traceability becomes a primary ethical requirement, ensuring that every cultural claim can be audited against authoritative source evidence. By applying the four pillars of RI\u0026mdash;anticipation, reflexivity, inclusion, and responsiveness\u0026mdash;the proposed framework moves beyond static evaluation. It ensures that the system's retrieval mechanisms are not only transparent but also responsive to the evolving ethical and social expectations of the source communities, thereby safeguarding the epistemic integrity of cultural knowledge through structured human oversight.\u003c/p\u003e \u003cp\u003eAs Tsamados et al. (\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e2022\u003c/span\u003e) highlight in their updated analysis of algorithmic ethics, alongside well-documented concerns about bias and fairness, issues of \"epistemic concerns\"\u0026mdash;including the justification of knowledge claims and the opacity of decision-making\u0026mdash;are central to the trustworthiness of AI systems (see also Mittelstadt et al., \u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e2016\u003c/span\u003e). Therefore, a framework for cultural chatbots must go beyond conventional evaluation to incorporate mechanisms for traceability and structured human oversight, ensuring that the system's epistemic authority is both accountable and aligned with the values of the communities whose heritage it mediates.\u003c/p\u003e \u003cp\u003eThis shift raises questions concerning epistemic accountability in algorithmic mediation rather than just factual accuracy.\u003c/p\u003e \u003cp\u003eCross-cultural communication theory emphasizes that the construction of meaning is mediated by culturally rooted interpretive frameworks (Ting-Toomey, \u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e2015\u003c/span\u003e). Cultural dimensions theory operationalizes variations in value orientations, power distance, and uncertainty tolerance across societies (Hofstede et al., \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e2010\u003c/span\u003e). In natural language processing research, cultural alignment is increasingly distinguished from linguistic fluency, necessitating context-sensitive adaptation mechanisms (Hershcovich et al., \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). Empirical studies have demonstrated measurable gaps between the outputs of generative models and the value distributions of cross-cultural societies (Cao et al., \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Lee et al., \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e2025\u003c/span\u003e). These findings suggest that generative systems may remain grammatically coherent while being epistemically misaligned with local cultural norms.\u003c/p\u003e \u003cp\u003eConsequently, tourism chatbots should not be evaluated solely as information delivery tools. They must be understood as culturally rooted speakers operating within normative and governance constraints. While the smart tourism literature acknowledges digital mediation, systematic mechanisms for assessing Cultural Faithfulness in generative tourism systems remain underdeveloped, particularly at the intersection of epistemic accountability and cultural representation.\u003c/p\u003e \u003cp\u003eConsidering this cultural context, we now examine the technical framework that facilitates these narratives through culturally rooted narrative agents and epistemically constrained generative processes.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2 Generative Models, Attribution Erosion, and Epistemically Constrained Retrieval-Augmented Generation\u003c/h2\u003e \u003cp\u003eThis section examines the technical characteristics of generative models within the cultural context established earlier. While these systems demonstrate significant creative capacity, they remain vulnerable to attribution erosion, unsubstantiated elaborations, and groundless fabrications (Ji et al., \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Bender et al., \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). These risks pose critical challenges in domains where narrative authority is culturally codified. Beyond surface-level inaccuracies, such models may generate seemingly coherent explanations that lack verifiable evidential foundations, reflecting a prioritization of performance over epistemic fidelity (Birhane et al., \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). Retrieval-Augmented Generation (RAG) mitigates these limitations through epistemic constraints that bind the generative process to external knowledge sources By anchoring outputs in retrieved evidence (Lewis et al., \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Guu et al., \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2020\u003c/span\u003e), subsequent improvements have refined retrieval conditioning, optimized knowledge integration, and enhanced scaling efficiency (Borgeaud et al., \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). Multi-hop architectures now enable contextual aggregation across distributed knowledge repositories (Asai et al., \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2022\u003c/span\u003e), while artifact-level traceability systems establish explicit links between generated content and source evidence segments (Tafjord et al., \u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). Explainable RAG frameworks further advance this approach by documenting intermediate reasoning pathways and retrieval decisions, thereby supporting auditability and transparency (Ren et al., \u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e2025\u003c/span\u003e). These methodological advances collectively reinforce narrative boundary stability and reduce ungrounded generation risks.\u003c/p\u003e \u003cp\u003eHowever, factual grounding does not inherently ensure epistemic accountability. Cross-cultural NLP research demonstrates that models may retrieve factually accurate information while subtly reframing narratives in culturally misaligned or normatively divergent ways (Hershcovich et al., \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Liu et al., \u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e2025\u003c/span\u003e). Surveys of evaluation frameworks reveal that benchmark-driven RAG assessments often prioritize lexical overlap and answer matching, frequently compromising contextual fidelity and narrative boundary stability (Gao et al., \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). Studies on factuality in summarization show that surface-level alignment with source texts does not guarantee cultural faithfulness, which systematically distinguishes between literal factual adherence and the preservation of culturally codified meaning (Maynez et al., \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e2020\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eStandard automated metrics such as ROUGE (Lin, \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2004\u003c/span\u003e) and BERTScore (Zhang et al., \u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e2020\u003c/span\u003e) measure semantic similarity but fail to detect plausible yet unsubstantiated reasoning, epistemic shifts, or culturally sensitive narrative distortions. Research on attribution erosion typologies indicates that generative errors exist on a spectrum from minor contextual deviations to severe fabrications, including epistemic misrepresentations that similarity-based metrics often overlook (Ji et al., \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). Furthermore, bias amplification and representational harms complicate evaluation, especially in culturally sensitive domains where narrative authority is culturally codified (Abid et al., \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). These distortions often arise when generative models fail to preserve the specific cultural and temporal nuances of word meanings, leading to narrative misalignment (Lucy \u0026amp; Bamman, \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). Retrieval-Augmented Generation (RAG) mitigates these limitations through epistemic constraints.\u003c/p\u003e \u003cp\u003eRecent studies highlight that improved retrieval quality does not inherently produce epistemically constrained generative outputs. Salemi and Zamani (\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e2024\u003c/span\u003e) demonstrate that systematic evaluation of retrieval components is essential for anchoring outputs in retrieved evidence within retrieval-augmented generation (RAG) systems. However, high retrieval performance does not ensure that generated narratives maintain contextual anchoring or cultural codification. Even with consistent retrieval accuracy, epistemically constrained generative processes may introduce interpretive shifts, subtle reframing, or contextual expansions. These findings emphasize the need for structured evaluation frameworks that assess both artifact-level traceability and the stability of cultural codification in generated narratives.\u003c/p\u003e \u003cp\u003eWhile RAG significantly enhances artifact-level traceability, current evaluation paradigms remain inadequate for detecting nuanced attribution erosion, culturally rooted reframing, and longitudinal shifts in narrative boundary stability.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3 Human Evaluation, Reliability, and Spectrum-Based Diagnosis\u003c/h2\u003e \u003cp\u003eAutomated metrics remain insufficient for capturing contextual nuances in culturally mediated AI systems, necessitating human evaluation as a critical validation mechanism. Inter-rater agreement metrics such as Cohen's kappa (Cohen, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e1960\u003c/span\u003e) and Fleiss' kappa (Fleiss, \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e1971\u003c/span\u003e) operationalize annotation consistency across evaluators, while computational linguistics research formalizes this annotation reliability within NLP evaluation pipelines (Artstein \u0026amp; Poesio, \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2008\u003c/span\u003e). Interpretive benchmarks (Landis \u0026amp; Koch, \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e1977\u003c/span\u003e; Krippendorff, \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e2019\u003c/span\u003e) establish structured thresholds for evaluating agreement strength, but in complex qualitative assessments involving extensive taxonomies, this rigor is further bolstered through negotiated agreement and intercoder reliability processes (Campbell et al., \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2013\u003c/span\u003e). These methodological foundations collectively ensure the systematic integrity required for evaluating complex categorical taxonomies in AI-generated content.\u003c/p\u003e \u003cp\u003eThe structured human adjudication paradigm integrates documented consensus adjudication into AI workflows, enabling iterative correction and supervision (Amershi et al., \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). From a human-centered AI perspective, oversight mechanisms must prioritize both system performance optimization and the enforcement of accountability through controlled deployment (Shneiderman, 2020). However, current implementations often emphasize model refinement and usability improvements over the systematic adjudication of narrative epistemic accountability in culturally sensitive domains.\u003c/p\u003e \u003cp\u003eCritiques of conventional human-in-the-loop models highlight the inadequacy of assuming human presence as a default guarantee for correct and ethical system function (Tschiatschek et al., \u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). A strategic transition is required, where the human role as a decision-maker is explicitly defined and supported by comprehensive information. This aligns with the concept of co-reasoning proposed by Salloch and Eriksen (\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e2024\u003c/span\u003e), wherein experts engage in collaborative evidence interpretation to achieve legitimate epistemic decisions through documented consensus.\u003c/p\u003e \u003cp\u003eStudies on dialectal and sociolinguistic variations demonstrate how subtle linguistic deviations can encode cultural distortions or marginalization (Blodgett et al., \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). Complementary research on attribution erosion typologies and epistemically constrained retrieval-augmented generation (RAG) indicates that generative errors are neither binary nor uniform; instead, they are distributed along a gradational continuum of attribution erosion (Ji et al., \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). The comprehensive survey by Ji et al. (\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2023\u003c/span\u003e) systematically catalogues the landscape of hallucination phenomena in natural language generation, providing a foundational framework for understanding how generative errors manifest across tasks and why evaluation must extend beyond surface-level metrics.\u003c/p\u003e \u003cp\u003eThese findings suggest that binary truth labels fail to capture the complexity of narrative misalignment. Collectively, this research trajectory justifies a spectrum-based diagnosis of attribution erosion. Rather than classifying outputs as simply true or false, a gradational scheme allows for the systematic differentiation between grounded responses, minor contextual shifts, moderate cultural misalignment, and severe fabrication. Such spectrum-based operationalization aligns with annotation reliability theory, bias taxonomy design, and structured human adjudication paradigms. Consequently, this approach provides the theoretical foundation for a multi-level diagnostic evaluation framework in culturally mediated generative systems, where narrative boundary stability is preserved through artifact-level traceability and epistemic accountability.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e2.4 Narrative Boundary Stability and Artifact-Level Traceability\u003c/h2\u003e \u003cp\u003eArtifact-level traceability systems enhance evidential transparency, yet the structural stability of narrative boundaries during epistemically constrained generative processes remains under-theorized. Generative systems conditioned on retrieved chunks frequently extrapolate beyond evidential limits, creating a hybridization of grounded content and unsubstantiated elaborations. This dynamic reflects an inherent conflict between probabilistic language modeling and epistemic constraints.\u003c/p\u003e \u003cp\u003eRetrieval-augmented architectures improve contextual anchoring through external knowledge integration during generation (Guu et al., \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Lewis et al., \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). Large-scale fine-tuning subsequently optimizes retrieval depth and parameter efficiency (Borgeaud et al., \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). Multi-hop reasoning systems further enhance cross-document aggregation and contextual linkage (Asai et al., \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). Nevertheless, increased retrieval capacity does not inherently prevent generative extrapolation beyond epistemic boundaries.\u003c/p\u003e \u003cp\u003eStudies of plausible reasoning errors show that models often produce semantically fluent yet unsubstantiated expansions (Ji et al., \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). This highlights a crucial distinction between artifact-level traceability and narrative boundary stability. Even with relevant artifacts retrieved, systems may generate extensions that subtly transgress evidential constraints while maintaining rhetorical coherence.\u003c/p\u003e \u003cp\u003eThe concept of boundary awareness, introduced within the symbiotic epistemology framework (Kapusta, \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e2025\u003c/span\u003e), offers a means to comprehend this phenomenon. Kapusta posits that a reliable epistemic agent is one capable of recognizing the limits of its knowledge and refraining from action when evidence is insufficient. The dual-level transparency principle in his TRACE protocol, encompassing high-level reasoning patterns and detailed factor explanations, provides a language for describing how boundary awareness can be operationalized.\u003c/p\u003e \u003cp\u003eFurthermore, Hickey (\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2026\u003c/span\u003e) argues that epistemic trust in opaque AI systems can be fostered through the provision of authoritative and verifiable artifact-level traceability systems. This approach diverges from efforts to demystify internal algorithmic black boxes. Within the RAG context, retrieved chunks serve as these constitutive knowledge sources. Technical explorations of chunk injection mechanisms for cultural tourism applications have demonstrated the feasibility of this approach, though systematic evaluation frameworks remain underdeveloped. By explicitly embedding chunks as contextual memory, artifact-level traceability becomes the bedrock for epistemic accountability.\u003c/p\u003e \u003cp\u003eNarrative boundary stability refers to the degree to which generated narratives remain confined within the epistemic scope defined by retrieved artifacts. Artifact-level traceability systems enhance traceability by explicitly linking outputs to ranges of supporting evidence (Tafjord et al., \u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). However, existing evaluation frameworks seldom assess how consistently generative outputs respect artifact-level traceability at the chunk level.\u003c/p\u003e \u003cp\u003eBy conceptualizing chunk injection as a contextual anchoring mechanism, narrative boundary stability emerges as an operationalizable construct linking artifact-level traceability with narrative containment. This framing extends the grounding literature by introducing artifact-level diagnostics that evaluate not only whether information is retrieved, but whether generative extrapolations remain epistemically constrained within the retrieved scope. Such artifact-sensitive evaluations advance the theoretical integration of retrieval conditioning and narrative governance in generative systems.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e2.5 Governance, Cultural Faithfulness, and Epistemic Accountability\u003c/h2\u003e \u003cp\u003eThe AI governance framework operationalizes transparency, accountability, fairness, and meaningful human control as foundational principles for responsible deployment (OECD, \u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Jobin et al., \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). The artifact-level traceability principle in the OECD AI Principles 2024 explicitly positions traceability as an accountability instrument to enable the analysis of outcomes and responses to inquiries (OECD, \u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). AI ethics scholarship further argues that governance frameworks must extend beyond abstract ethical principles to incorporate culturally codified epistemic accountability, particularly in contexts where knowledge production intersects with identity and heritage (Mohamed, Png, \u0026amp; Isaac, \u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e2020\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eWithin tourism mediation settings, tensions between algorithmic storytelling and indigenous interpretive authority have been empirically documented. Previous studies examining automated systems in cultural contexts have demonstrated how such systems can inadvertently reshape cultural representations through generative processes (Nurhidayati et al., \u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e2025\u003c/span\u003e). However, technical efforts to enhance retrieval reliability in tourism chatbots suggest that epistemic accountability is operationally feasible. A broader epistemic analysis of generative tourism systems reveals persistent risks of narrative compression, oversimplification, and cultural distortion despite retrieval augmentation (Nurhidayati et al., \u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e2026\u003c/span\u003e). These findings indicate that governance frameworks must extend beyond high-level ethical principles to implement structured human adjudication paradigms that preserve narrative boundary stability through artifact-level traceability systems.\u003c/p\u003e \u003cp\u003eOperational epistemic accountability necessitates the operationalization of governance principles into reproducible evaluative procedures. This requires severity-differentiating diagnostics, structured human adjudication protocols, and reproducible artifact preservation to enable cross-version comparison and longitudinal auditability. The principles of reproducibility and artifact-level traceability in computational research increasingly underscore the need to document methodological evolution and maintain evidential outputs for verification.\u003c/p\u003e \u003cp\u003eIn culturally sensitive domains, governance becomes inseparable from evaluation architecture. Artifact-level traceability systems, severity-differentiating diagnostics for epistemic deviations, and preserved diagnostic artifacts constitute concrete instantiations of governance principles. The AI4People framework proposed by Floridi and colleagues (\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2018\u003c/span\u003e) provides an ethical foundation with the principles of beneficence, non-maleficence, autonomy, justice, and epistemic accountability. The principle of epistemic accountability, in particular, demands that AI systems be explainable and answerable. This demand is met by the artifact-level traceability systems within our framework.\u003c/p\u003e \u003cp\u003eFurthermore, Laux and colleagues' (\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e2024\u003c/span\u003e) critique of equating trustworthiness with risk acceptability in the EU AI Act highlights the need for a richer definition of trust in AI. Our framework contributes to this discourse by operationalizing Cultural Faithfulness as a crucial component of trustworthiness within the context of cultural tourism. The Human-Centered AI vision by Shneiderman (\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e2022\u003c/span\u003e) emphasizes the importance of human control and transparent reporting. This vision is embodied in our framework's design. The principles of artifact-level traceability and structured human adjudication that we implement are concrete manifestations of that vision for cultural domains. By embedding governance into evaluation design, cultural protection shifts from declarative policy aspirations to operational accountability.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e2.6 Research Gaps and Integrative Positioning\u003c/h2\u003e \u003cp\u003eThe reviewed literature reveals fragmentation across four interconnected domains: digital tourism communication, retrieval engineering, human evaluation methodologies, and AI governance. Smart tourism research acknowledges AI-mediated storytelling and digital interactivity (Gretzel et al., \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2015\u003c/span\u003e; Sigala, \u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e2018\u003c/span\u003e). However, the systematic operationalization of Cultural Faithfulness as an evaluative construct remains limited. Retrieval-Augmented Generation (RAG) architectures enhance factual grounding through contextual anchoring (Lewis et al., \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Gao et al., \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). Yet, evaluation protocols predominantly emphasize lexical similarity and benchmark optimization (Lin, \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2004\u003c/span\u003e; Zhang et al., \u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). Reliability theory formalizes inter-rater consensus (Artstein \u0026amp; Poesio, \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2008\u003c/span\u003e), although structured epistemic diagnostics are seldom integrated into the evaluation of applied generative systems. Concurrently, governance frameworks articulate principles of transparency and accountability (OECD, \u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Jobin et al., \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). These norms are rarely operationalized through artifact-level evaluation protocols.\u003c/p\u003e \u003cp\u003eA critical limitation in current scholarship lies in the absence of a unified evaluative architecture that systematically integrates artifact-level traceability systems, severity-differentiating diagnostics, structured human adjudication protocols, and reproducible artifact preservation mechanisms. While prior research has examined these elements in isolation, their coordinated implementation remains insufficiently developed, especially in culturally sensitive domains such as tourism.\u003c/p\u003e \u003cp\u003eThis research proposes a traceable structured human adjudication framework for epistemically constrained generative processes in tourism chatbots. The framework combines chunk injection as contextual anchoring mechanisms, severity-differentiating diagnostics, consensus-based reliability validation, and artifact-level traceability systems. By operationalizing Cultural Faithfulness through epistemic accountability principles, this work establishes an interdisciplinary nexus between smart tourism research and responsible AI evaluation, ensuring narrative boundary stability through artifact-level traceability and epistemic constraint enforcement.\u003c/p\u003e \u003c/div\u003e"},{"header":"3. Methodology","content":"\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Research Design and Analytical Scope\u003c/h2\u003e \u003cp\u003eThis study employs a structured evaluative design grounded in the structured human adjudication paradigm. Its primary objective extends beyond measuring semantic similarity to examining whether retrieval functions as epistemic constraints on epistemically constrained generative processes within culturally codified tourism narratives.\u003c/p\u003e \u003cp\u003eThe analysis unit comprises 61 system responses corresponding to 61 evaluation questions. Each response undergoes two complementary analytical procedures: automated similarity screening and structured human adjudication using severity-differentiating diagnostics.\u003c/p\u003e \u003cp\u003eTwo levels of analysis are defined. At the question level, 61 final consensus decisions are examined to determine grounding classification patterns. At the annotation level, 183 independent scores generated by three evaluators are analyzed to assess diagnostic distribution and agreement structure. This dual structure ensures the evaluation captures both output categorization and more granular epistemic deviation patterns.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Dataset Construction and Question Design\u003c/h2\u003e \u003cp\u003eThe evaluation dataset (Version 1.2) was constructed from curated cultural documents sourced from the Ande-Ande Lumut narrative (see Supplementary Material A for the complete dataset). As discussed in the Theoretical Foundation, Cultural Faithfulness in RAG systems requires epistemically constrained references. Consequently, the narrative corpus functions as a controlled ground-truth knowledge base rather than a general domain benchmark.\u003c/p\u003e \u003cp\u003ePrior to retrieval indexing, source documents were segmented into semantically coherent text units. Each segment was recorded in the chunk segmentation table containing unique identifiers, exact text ranges, and their positions within the source documents. This segmentation enables artifact-level traceability and direct mapping between generated claims and evidential ranges.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDistribution of Question Types in Dataset v1.2 (n\u0026thinsp;=\u0026thinsp;61)\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eQuestion Type\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNumber\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAnalytical Purpose\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eRelated RQ\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eIn-context\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e30\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMeasure factual grounding accuracy and evidence alignment\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eRQ1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAmbiguous\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTest contextual sensitivity and interpretative restraint\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eRQ1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOut-of-context\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e15\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTest\u0026nbsp;attribution erosion\u0026nbsp;and boundary control\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eRQ1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTotal\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e61\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eStructured diagnostic evaluation dataset\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eTable source: Research data (can be verified in Supplementary Material A)\u003c/em\u003e \u003c/p\u003e \u003cp\u003eAs shown in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, the dataset comprises 30 in-context questions, 16 ambiguous questions, and 15 out-of-context questions. In-context questions assess the accuracy of grounding in the available evidence. Ambiguous questions evaluate contextual sensitivity and interpretative restraint. Out-of-context questions test attribution erosion behavior and resistance to unsubstantiated expansions. These three categories operationalize the attribution erosion typologies identified in the Theoretical Foundation.\u003c/p\u003e \u003cp\u003eTo test the stability of narrative boundaries under repeated generations, a subset of 14 questions was administered again under controlled conditions.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDistribution of Repeated Questions (n\u0026thinsp;=\u0026thinsp;14)\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eQuestion Category\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNumber Repeated\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAnalytical Purpose\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eRelated RQ\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eIn-context\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTest grounding stability under repetition\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eRQ4\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAmbiguous\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTest interpretative consistency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eRQ4\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOut-of-context\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAssess\u0026nbsp;abstention mechanism\u0026nbsp;reliability\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eRQ4\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTotal Repeated\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNarrative boundary stability\u0026nbsp;analysis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eTable source: Research data (can be verified in Supplementary Material A)\u003c/em\u003e \u003c/p\u003e \u003cp\u003eRepeated questions enable the assessment of artifact-level traceability and epistemic constraint enforcement across diverse generations.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e\u003cb\u003e3.3 Retrieval Architecture with Chunk Injection and Structured Human Adjudication\u003c/b\u003e.\u003c/h2\u003e \u003cp\u003eThe proposed evaluation architecture integrates retrieval, generation, filtering, and structured human adjudication into a unified workflow, as illustrated in Fig.\u0026nbsp;1.\u003c/p\u003e \u003cp\u003e[ Insert Fig.\u0026nbsp;1 here ]\u003c/p\u003e \u003cp\u003e \u003cb\u003eFigure 1\u003c/b\u003e. Structured Human Adjudication Architecture for Cultural Faithfulness, Showing Retrieval-Augmented Generation (RAG) Components\u003c/p\u003e \u003cp\u003eThe proposed framework integrates four core methodological components: (1) artifact-level traceability through chunk injection, (2) severity-differentiating diagnostics using the E0\u0026ndash;E10 taxonomy, (3) structured human adjudication with documented consensus procedures, and (4) reproducible artifact preservation.\u003c/p\u003e \u003cp\u003eThe architecture operates as follows. Culturally codified knowledge sources ① are segmented into indexed text chunks ② and explicitly injected into model context ③, establishing artifact-level traceability. A single query triggers two parallel retrieval-augmented generation (RAG) paths. The RAG system (Candidate Generator) ④ produces candidate responses, while the RAG system (Reference \u0026amp; Chunk_ID Generator) ⑤ generates reference responses accompanied by supporting chunk identifiers. This dual-path configuration enables direct mapping between generated claims and source evidence.\u003c/p\u003e \u003cp\u003eAn automated metric filtering layer ⑥ computes similarity scores, including BERTScore F1, ROUGE-L F1, and cosine similarity, between candidate and reference responses. These metrics function solely as analytical signals and do not incorporate chunk identifiers, thereby avoiding their use as grounding validators. Threshold-based similarity labeling provides early indicators of potential epistemic deviation. In the dataset, these values are recorded as bertscore_f1, rougeL_f1, and cosine_similarity.\u003c/p\u003e \u003cp\u003eSubsequently, three evaluators independently perform structured human adjudication ⑦ using the E0\u0026ndash;E10 severity-differentiated taxonomy. Evaluators are granted full access to source documents, chunk segmentation tables, retrieved chunk identifiers, candidate and reference responses, and automated filtering outputs. Cases meeting predefined trigger conditions, such as evaluator divergence or high-severity classifications, undergo documented consensus adjudication ⑧ to achieve unanimous final decisions. Individual annotations are stored with evaluator_id, error_code, and annotation_notes fields.\u003c/p\u003e \u003cp\u003eAll artifacts generated throughout the pipeline, including source documents, segmentation tables, similarity metrics, and consensus resolutions, are systematically archived under reproducible artifact preservation ⑨.\u003c/p\u003e \u003cp\u003eThe archive includes 183 independent annotations, 61 final decisions, evaluation questions, generated responses, chunk mappings, filtering outputs, and consensus metadata. The complete anonymized archive is provided as Supplementary Material A.\u003c/p\u003e \u003cp\u003eThis architecture operationalizes narrative boundary stability by verifying whether generated outputs remain confined within epistemically constrained evidential ranges during generation. The structured adjudication layer embeds interpretative validation directly within the retrieval workflow, addressing the evaluative integration gap identified in the Theoretical Foundation.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003e3.4 Automated Similarity Filtering\u003c/h2\u003e \u003cp\u003ePost-generation, candidate responses and reference responses underwent automated similarity analysis. Similarity metrics were computed using BERTScore F1, ROUGE-L F1, and cosine similarity. Threshold-based similarity classifications and early indicators of potential grounding issues were generated to support interpretive analysis. In the dataset, these are recorded as similarity_label (high/medium/low) and early_warning_flags.\u003c/p\u003e \u003cp\u003eNotably, similarity calculations were restricted to pairwise comparisons between candidate and reference responses. Chunk identifiers were excluded from metric computations, ensuring that filtering functioned as an analytical signal rather than a definitive grounding validator.\u003c/p\u003e \u003cp\u003eAutomated metrics did not dictate final evaluation outcomes. Instead, these metrics provided structured input for documented consensus adjudication. All metric outputs were preserved within the evaluation archive (Supplementary Material A).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003e3.5 Granular Cultural Faithfulness Scale (E0\u0026ndash;E10)\u003c/h2\u003e \u003cp\u003eNarrative boundary stability is evaluated through an 11-point diagnostic framework spanning E0 (fully grounded) to E10 (system failure). This severity-differentiating taxonomy surpasses binary detection approaches by identifying nuanced degrees of epistemic constraint violation and unsubstantiated expansions.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eGranular Cultural Faithfulness Scale (E0\u0026ndash;E10)\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCode\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCategory\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDescription\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eEpistemic Severity\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eFailure Type\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eE0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFully Grounded\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eResponse is fully relevant, correctly answers the query, and is explicitly supported by the referenced chunk.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eNone\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eE1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMinor Omission\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eResponse is correct and source-supported, but omits minor details that do not alter the main meaning.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eVery Low\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eContent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eE2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePartial Grounding\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eResponse is partially supported by the chunk but includes additional information not fully traceable to the source.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eLow\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eContent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eE3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eWeak Attribution\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eResponse is relevant, but the referenced chunk supports only a limited portion of the content.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eModerate\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eContent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eE4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eContext Drift\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eResponse deviates from the focus of the query while remaining within the same general domain.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eModerate\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eContent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eE5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAmbiguous Response\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eResponse is overly general or ambiguous, preventing clear verification against the source.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eElevated\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eContent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eE6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eUnsupported Claim\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eResponse contains factual claims not supported by any chunk.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eHigh\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eContent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eE7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAttribution Erosion\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eResponse appears linguistically plausible but lacks valid source support in the chunks.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSevere\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eContent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eE8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eContradictory Content\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eResponse directly contradicts information contained in the source chunk.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSevere\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eContent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eE9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAbstain / No-Answer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eResponse explicitly declines to answer or avoids providing information.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eNon-error (controlled)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eE10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSystem Failure\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eSystem fails to produce an evaluable output, such as empty output or processing error.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eCritical\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eSystem\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eTable source: Developed from the proposed evaluation framework and validated through annotation. Raw annotation data are available in Supplementary Material A.\u003c/em\u003e \u003c/p\u003e \u003cp\u003eImportant notes for alignment with datasets:\u003c/p\u003e \u003cp\u003eThe annotation scheme implemented in this study used the category codes E0\u0026ndash;E10 as documented in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e. In the dataset (Supplementary Material A), these codes are stored in the error_code field. It should be noted that category E7, originally labeled as \"Hallucination\" during the initial annotation phase, has been conceptually refined to \"Attribution Erosion\" based on post-hoc analysis of error patterns. This refinement captures the observation that generative errors more frequently manifest as gradual erosion of source attribution rather than complete fabrication. The raw annotation data for E7 remains unchanged and verifiable in the dataset. Similarly, E9 (\"Abstain\") is conceptualized as a \"calibrated non-response\" and a positive indicator of boundary awareness, rather than a system failure.\u003c/p\u003e \u003cp\u003eAs presented in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e, this taxonomy categorizes responses across eleven distinct classes: fully supported claims, paraphrastic variations, interpretive extensions, contextual shifts, weak attributions, unsubstantiated expansions, attribution erosion, contradictions, mixed grounding failures, deliberate omissions (abstention mechanism), and system errors. The analytical separation of abstentions from attribution erosion and system errors enables evaluation of calibrated non-responses as epistemic accountability indicators rather than failure metrics.\u003c/p\u003e \u003cp\u003eBefore formal assessment, a calibration session established inter-rater consensus on interpretive thresholds without pre-determined outcome expectations.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003e3.6 Evaluator Selection and Structured Human Adjudication\u003c/h2\u003e \u003cp\u003eThree evaluators were selected based on their domain expertise and competencies in culturally codified meaning interpretation. All evaluators held graduate-level qualifications in cultural studies or linguistics, with specific expertise in Indonesian folklore and narrative traditions. One evaluator additionally demonstrated proficiency in natural language processing and retrieval systems.\u003c/p\u003e \u003cp\u003eTraining encompassed operationalization of the E0\u0026ndash;E10 taxonomy, structured artifact alignment exercises, and consensus adjudication pilots. The evaluation of 61 responses yielded 183 independent annotations, with each response receiving three independent scores. In the dataset, these are recorded with unique annotation_id values linked to question_id and evaluator_id.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003e3.7 Selective Consensus Adjudication\u003c/h2\u003e \u003cp\u003eFollowing independent assessments, evaluators conducted structured human adjudication sessions. Cases were flagged for deliberation when attribution erosion typologies (E7\u0026ndash;E10) emerged, substantial divergences among evaluators occurred, or potential grounding issues were identified.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eConsensus Trigger Conditions and Resolution Structure\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTrigger Condition\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDescription\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAction Taken\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eResolution Requirement\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRelated RQ\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAttribution Erosion Typologies (E7\u0026ndash;E10)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAny annotation E7 to E10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eCase flagged for review\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eUnanimous agreement required\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRQ2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eScore Divergence\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSeverity level difference\u0026thinsp;\u0026ge;\u0026thinsp;3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eStructured deliberation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eArtifact re-examination with full traceability\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRQ2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAttribution Erosion Indicators\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePresence of E4 classification or ambiguous cases\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTargeted chunk verification\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eEvidence confirmation through source mapping\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRQ2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eConvergent Low Severity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAll scores within E0\u0026ndash;E3 range\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo extended deliberation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMajority convergence accepted\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRQ2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCalibrated Non-Response\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eE9 assigned\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eConfirm deliberate omission\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eEvidence absence verification\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRQ2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eTable source: Documented consensus adjudication protocol developed for this study.\u003c/em\u003e \u003c/p\u003e \u003cp\u003eMarked cases undergo collective artifact re-examination, including ground-truth documents, chunk segmentation tables, retrieved chunk identifiers, candidate responses, reference responses, and automated filtering outputs. Final decisions require unanimous agreement. In the dataset, these deliberations are documented in consensus_notes and the final decision is recorded in final_decision (accept/reject) fields.\u003c/p\u003e \u003cp\u003eUnmarked cases are finalized based on convergence without extended deliberation. This selective mechanism operationalizes human judgment as a structured epistemic accountability framework rather than a post-hoc validation layer. This approach directly addresses RQ2 and the identified evaluative integration gaps in the Theoretical Foundation.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003e3.8 Traceable Artifact Preservation and Reproducible Design\u003c/h2\u003e \u003cp\u003eAll evaluation components are archived and available as Supplementary Material A. To maintain anonymity during double-blind peer review, the dataset is presented with a generic title (\"Supplementary Dataset for Cultural Tourism Chatbot Evaluation\") and without author identifiers. The file is under embargo during review and will be made publicly available upon publication.\u003c/p\u003e \u003cp\u003eThe archive contains:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e183 independent annotations (with evaluator_id, error_code, annotation_notes)\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e61 final consensus decisions (final_decision, consensus_notes)\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eCurated source documents and chunk segmentation tables\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e61 evaluation questions with category labels (in-context/ambiguous/out-of-context)\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eCandidate and reference responses for all questions\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eRetrieved chunk identifiers for each response\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eAutomated filtering outputs (similarity scores, labels)\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eEach evaluation decision is traceable to specific textual evidence and documented deliberative rationale. By preserving annotation-level and decision-level records, this framework operationalizes auditability, reproducible artifact preservation, and epistemic accountability aligned with governance standards. This approach directly addresses RQ5 and translates transparency and accountability principles into procedural implementation.\u003c/p\u003e \u003c/div\u003e"},{"header":"4. Results","content":"\u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003e4.1 RQ1: Measurable Levels of Cultural Faithfulness\u003c/h2\u003e \u003cp\u003eThe first research question examined the discriminative capacity of tourism chatbots utilizing epistemically constrained Retrieval-Augmented Generation architectures with chunk injection as contextual anchoring mechanisms to demonstrate Cultural Faithfulness under structured granular evaluation.\u003c/p\u003e \u003cp\u003eOf the 61 evaluated questions, 45 responses (73.8%) were accepted while 16 responses (26.2%) were rejected following documented consensus adjudication. The final decision distribution is presented in Table\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eFinal Consensus-Based Outcomes (n\u0026thinsp;=\u0026thinsp;61)\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFinal Decision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFrequency\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePercentage (%)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAccepted\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e73.8\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRejected\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e26.2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTotal\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e61\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e100\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eSource: Research data (can be verified in Supplementary Material A)\u003c/em\u003e \u003c/p\u003e \u003cp\u003eAccepted responses included outputs categorized by grounding severity levels (E0 to E2) and calibrated non-responses (E9) in out-of-context scenarios. Rejected responses encompassed attribution erosion (E7), contradictions (E8), system failures (E10), and severe unsubstantiated elaborations. The coexistence of accepted and rejected cases indicates that the evaluation framework does not artificially inflate performance. Instead, this framework operationalizes a measurable differentiation between epistemically grounded cultural representations and invalid narrative outputs.\u003c/p\u003e \u003cp\u003eCultural Faithfulness is thus conceptualized not as a binary attribute but as a structured evaluative construct integrating artifact-level traceability, interpretive restraint, and attribution transparency. These findings demonstrate that in culturally codified tourism corpora, the combination of chunk-level retrieval and granular assessment enables observable and measurable narrative boundary stability preservation.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003e4.2 RQ2: Contribution of the Structured Human Adjudication Paradigm\u003c/h2\u003e \u003cp\u003eThe second research question examines how the structured human adjudication paradigm complements artifact-level traceability systems in enhancing epistemic accountability beyond automated similarity metrics.\u003c/p\u003e \u003cp\u003eOf the 61 evaluated cases, 16 (26.2%) were flagged for documented consensus adjudication prior to final evaluation, while 45 (73.8%) proceeded without escalation. This distribution is presented in Table\u0026nbsp;\u003cspan refid=\"Tab6\" class=\"InternalRef\"\u003e6\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab6\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 6\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eFlagged vs Non-Flagged Cases (n\u0026thinsp;=\u0026thinsp;61)\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCategory\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFrequency\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePercentage (%)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFlagged for Documented Consensus Adjudication\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e26.2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNot Flagged\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e73.8\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTotal\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e61\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e100\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eSource: Research data (can be verified in Supplementary Material A)\u003c/em\u003e \u003c/p\u003e \u003cp\u003eAll sixteen flagged cases were rejected following artifact-level re-examination. In contrast, forty-five non-flagged cases were accepted without extended deliberation. The direct alignment between flagging status and final decisions demonstrates that structured escalation protocols effectively identified high-risk responses prior to structured human adjudication.\u003c/p\u003e \u003cp\u003eAutomated similarity metrics did not influence final outcomes. Flagged cases underwent collaborative artifact review involving curated source documents, chunk segmentation tables, retrieved chunk identifiers, and candidate-reference response comparisons. Final decisions were determined exclusively through documented consensus adjudication.\u003c/p\u003e \u003cp\u003eThese findings reveal that the structured human adjudication paradigm operates as an epistemic accountability mechanism rather than a procedural addendum. Within culturally codified narrative contexts, interpretative discrepancies may emerge despite high lexical similarity. The structured human adjudication protocols therefore function as interpretative validation frameworks, addressing subtle narrative boundary instabilities imperceptible to automated metrics.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section2\"\u003e \u003ch2\u003e4.3 RQ3: Distribution Patterns and Epistemic Accountability Calibration\u003c/h2\u003e \u003cp\u003eThe third research question examines whether the granular scale from E0 to E10 generates meaningful distribution patterns and reflects calibrated epistemic accountability. From 183 independent annotations (61 questions \u0026times; 3 evaluators), the granular score distribution is presented in Table\u0026nbsp;\u003cspan refid=\"Tab7\" class=\"InternalRef\"\u003e7\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab7\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 7\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDistribution of Granular Scores (E0 to E10) (n\u0026thinsp;=\u0026thinsp;183)\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eScore\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFrequency\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePercentage (%)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eE0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e57\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e31.1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eE1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e3.3\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eE2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e39\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e21.3\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eE3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e2.7\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eE4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e27\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14.8\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eE5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eE6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eE7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.6\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eE8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.6\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eE9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e26\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14.2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eE10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e15\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e8.2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTotal\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e183\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e100\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eSource: Research data (can be verified in Supplementary Material A)\u003c/em\u003e \u003c/p\u003e \u003cp\u003eMost annotations cluster within the E0 to E4 range, indicating responses predominantly anchored in source evidence with limited interpretative extensions. High-severity categories (E7 and E8) appear exclusively in rejected cases and remain infrequent, demonstrating that attribution erosion and contradictions are controlled rather than pervasive.\u003c/p\u003e \u003cp\u003eE9 (abstention mechanism) emerges in both accepted and rejected contexts. Accepted E9 instances align with appropriate epistemic restraint for out-of-context queries, whereas rejected E9 instances reflect non-substantive non-responses or procedural insufficiencies. This distinction confirms that abstention functions as a calibrated epistemic accountability mechanism rather than a performance failure.\u003c/p\u003e \u003cp\u003eThe absence of E5 assignments reveals that ambiguous moderate attribution weaknesses are not characteristic of evaluated outputs. Annotations cluster toward clearly grounded or evidently invalid categories, indicating effective taxonomic discriminative capacity. Collectively, this distribution demonstrates layered severity differentiation and supports the claim that the E0\u0026ndash;E10 scale enables structured epistemic accountability calibration beyond binary loyal/disloyal classifications.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec22\" class=\"Section2\"\u003e \u003ch2\u003e4.4 RQ4: Role of Chunk-Injection in Narrative Boundary Stability\u003c/h2\u003e \u003cp\u003eThe fourth research question investigates the contribution of the retrieval architecture with chunk injection to narrative consistency and verifiable grounding across the evaluated dataset. The results show that none of the accepted responses contained attribution erosion or contradictions, and all accepted responses maintained traceable alignment between generated claims and identified chunk segments within the curated corpus. This suggests that chunk-level retrieval controls inferential extension within document-bound narrative boundaries, and generative variability does not lead to catastrophic epistemic deviations in accepted outputs.\u003c/p\u003e \u003cp\u003eTo evaluate architectural consistency under identical semantic conditions, fourteen base questions were repeated, producing fourteen repeated responses (total 28 responses). The evaluator scores for each pair are presented in Table\u0026nbsp;8. For analytical clarity, severity ranges were categorized as follows:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eE0\u0026ndash;E4: Low-severity grounded range\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eE5\u0026ndash;E6: Moderate-severity range\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eE7\u0026ndash;E8: Attribution erosion or contradiction range\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eE9\u0026ndash;E10: Abstention mechanism or system response range\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003e\u003cstrong\u003eTable 8.\u003c/strong\u003e Evaluator Scores, Severity Range, and Final Decisions for Repeated Question Pairs (n = 14 Pairs)\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eQuestion Pair\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eInitial Scores (E0\u0026ndash;E10)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eInitial Severity Range\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eFinal Decision (Initial)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eRepeated Scores (E0\u0026ndash;E10)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eRepeated Severity Range\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eFinal Decision (Repeated)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eStability\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eQ0\u0026ndash;Q15\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE1, E1, E2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE0\u0026ndash;E4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAccepted\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE1, E2, E2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE0\u0026ndash;E4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAccepted\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eStable\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eQ02\u0026ndash;Q16\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE2, E3, E2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE0\u0026ndash;E4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAccepted\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE3, E2, E3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE0\u0026ndash;E4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAccepted\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eStable\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eQ03\u0026ndash;Q17\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE4, E4, E3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE0\u0026ndash;E4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAccepted\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE4, E4, E4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE0\u0026ndash;E4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAccepted\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eStable\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eQ04\u0026ndash;Q18\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE0, E0, E1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE0\u0026ndash;E4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAccepted\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE0, E1, E0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE0\u0026ndash;E4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAccepted\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eStable\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eQ05\u0026ndash;Q19\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE2, E3, E2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE0\u0026ndash;E4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAccepted\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE3, E4, E3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE0\u0026ndash;E4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAccepted\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eStable\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eQ06\u0026ndash;Q20\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE4, E4, E4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE0\u0026ndash;E4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAccepted\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE4, E3, E4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE0\u0026ndash;E4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAccepted\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eStable\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eQ07\u0026ndash;Q21\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE9, E9, E9\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE9\u0026ndash;E10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAccepted\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE9, E9, E9\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE9\u0026ndash;E10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAccepted\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eStable\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eQ08\u0026ndash;Q22\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE9, E9, E10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE9\u0026ndash;E10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAccepted*\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE9, E9, E9\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE9\u0026ndash;E10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAccepted\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eStable\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eQ09\u0026ndash;Q23\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE2, E1, E2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE0\u0026ndash;E4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAccepted\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE3, E4, E3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE0\u0026ndash;E4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAccepted\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eStable\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eQ10\u0026ndash;Q24\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE3, E4, E3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE0\u0026ndash;E4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAccepted\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE6, E6, E6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE5\u0026ndash;E6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eRejected\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eVariation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eQ11\u0026ndash;Q25\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE4, E4, E3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE0\u0026ndash;E4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAccepted\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE6, E6, E7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE5\u0026ndash;E8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eRejected\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eVariation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eQ12\u0026ndash;Q26\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE6, E6, E6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE5\u0026ndash;E6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eRejected\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE4, E4, E3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE0\u0026ndash;E4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAccepted\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eVariation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eQ13\u0026ndash;Q27\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE9, E9, E9\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE9\u0026ndash;E10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAccepted\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE2, E3, E2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eE0\u0026ndash;E4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAccepted\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eVariation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eQ13\u0026ndash;Q28\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eE6, E9, E10\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eE5\u0026ndash;E10\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eRejected\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eE9, E9, E9\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eE9\u0026ndash;E10\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eAccepted\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eVariation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cem\u003eNote: Q08 initial had one E10 (system failure) but was accepted after consensus as E10 was due to temporary system error; repeated generation resolved the issue.\u003c/em\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eSource: Reconstruction from original research data (verifiable in Supplementary Material A)\u003c/em\u003e\u003c/p\u003e \u003cp\u003eThe pair Q13\u0026ndash;Q28 (highlighted) demonstrates the stochastic nature inherent in large language models. In the initial generation (Q13), divergent evaluator scores (E6, E9, E10) indicated a complex scenario where the system generated an unsupported claim while the reference generation failed. After documented consensus adjudication, the response was rejected. In the repeated generation (Q28), both candidate and reference responses appropriately abstained, leading to acceptance. Notably, this variation did not escalate into attribution erosion (E7) or contradiction (E8) categories, showing that although generative outputs may differ, they remain constrained within epistemically controlled boundaries.\u003c/p\u003e \u003cp\u003eOf the fourteen tested pairs:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e10 pairs (71.4%) remained within the same severity range\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e4 pairs (28.6%) exhibited limited cross-range variation\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eNo paired instances escalated to attribution erosion (E7) or contradiction (E8)\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eNo contradictory chunk attributions were detected\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eThese results highlight the necessity of iterative testing and systematic human evaluation when assessing culturally sensitive generative systems. Consistency in this study is defined not as identical evaluator judgments, but as the preservation of severity structure and grounding alignment under repeated semantic inputs. These results indicate that chunk-level retrieval contributes to the preservation of narrative boundaries and containment of severity in culturally rooted generative tasks.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec23\" class=\"Section2\"\u003e \u003ch2\u003e4.5 RQ5: Traceability, Auditability, and Reproducibility\u003c/h2\u003e \u003cp\u003eThe fifth research question evaluates the traceability and governance alignment of the evaluation process. The study maintained comprehensive links for all 61 questions, including candidate and reference responses, chunk identifiers, filtering outputs, evaluator scores, annotation statuses, consensus documentation, and final decisions.\u003c/p\u003e \u003cp\u003eBoth rejected and accepted cases preserved annotation transparency, allowing for the reproduction of evaluation outcomes through explicit procedural links. Each evaluation decision can be traced to specific textual evidence and documented deliberative rationale via artifact-level traceability systems.\u003c/p\u003e \u003cp\u003eThe complete evaluation archive, documented in Supplementary Material A, contains:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e183 independent annotations with evaluator identifiers and error codes\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e61 final consensus decisions with resolution metadata\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eCurated source documents and chunk segmentation tables\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eAll candidate and reference responses with associated chunk identifiers\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eAutomated filtering outputs (similarity scores and labels)\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eThis structured documentation operationalizes governance-aligned evaluations in generative systems, which is essential in culturally sensitive tourism contexts where narrative legitimacy and epistemic accountability intersect. By preserving annotation-level and decision-level records, this framework demonstrates how abstract transparency and accountability principles can be translated into concrete, verifiable procedures.\u003c/p\u003e \u003c/div\u003e"},{"header":"5. Discussion","content":"\u003cp\u003eThis section interprets the findings presented in Section \u003cspan refid=\"Sec18\" class=\"InternalRef\"\u003e4\u003c/span\u003e in relation to the research questions and the theoretical framework established earlier. The discussion connects empirical results to prior literature, highlighting how the proposed evaluation framework advances the operationalization of Cultural Faithfulness through artifact-level traceability, severity-differentiating diagnostics, and structured human adjudication. The section is organized around the five research questions, followed by a discussion of limitations and future research directions.\u003c/p\u003e \u003cdiv id=\"Sec25\" class=\"Section2\"\u003e \u003ch2\u003e5.1 RQ1: Discriminative Capacity of the Evaluation Framework\u003c/h2\u003e \u003cp\u003eThe first research question examined the discriminative capacity of the proposed framework in distinguishing between grounded responses, attribution erosion, and severe grounding violations. The finding that 73.8% of responses were accepted while 26.2% were rejected confirms that the intentionally constructed dataset comprising in-context, ambiguous, and out-of-context questions successfully created challenging evaluation conditions. Rejections occurred exclusively in cases with significant grounding weaknesses, demonstrating that Cultural Faithfulness is not a binary attribute but a hierarchical construct requiring evaluation instruments sensitive to interpretive nuances.\u003c/p\u003e \u003cp\u003eThis result reinforces earlier cross-cultural NLP research showing that AI systems may produce grammatically coherent outputs that are culturally misaligned (Cao et al., \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Lee et al., \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e2025\u003c/span\u003e). The E0\u0026ndash;E10 taxonomy enabled finer differentiation beyond factual accuracy, supporting the argument that cultural representation in generative systems cannot be adequately assessed through surface-level metrics alone. The dataset's success in triggering severity score variation suggests that scenario-based approaches (in-context, ambiguous, out-of-context) can serve as a model for Cultural Faithfulness evaluation in other domains.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec26\" class=\"Section2\"\u003e \u003ch2\u003e5.2 RQ2: Role of Structured Human Adjudication\u003c/h2\u003e \u003cp\u003eThe second research question investigated how structured human adjudication complements automated similarity filtering. The finding that 16 cases (26.2%) were flagged for consensus review, all of which were rejected after artifact-level re-examination, demonstrates that automated metrics alone are insufficient for detecting subtle attribution erosion and unsubstantiated expansions. Even responses with high lexical similarity scores were rejected when human evaluators identified grounding failures.\u003c/p\u003e \u003cp\u003eThese findings align with critiques of conventional human-in-the-loop models (Tschiatschek et al., \u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e2024\u003c/span\u003e), which argue that human involvement is often superficial due to insufficient informational support. In our framework, evaluators were not merely final reviewers but decision-makers supported by rich artifacts: source chunks, traceability identifiers, similarity metrics, and evaluation records. The collective deliberation on flagged cases reflects what Salloch and Eriksen (\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e2024\u003c/span\u003e) term \u003cem\u003eco-reasoning\u003c/em\u003e, where experts jointly interpret evidence to achieve legitimate epistemic consensus.\u003c/p\u003e \u003cp\u003eThus, automated metrics function appropriately as an initial analytical layer that enriches human judgment rather than replacing it. This architecture operationalizes the procedural transparency demanded by AI governance principles (OECD, \u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e2024\u003c/span\u003e), positioning structured human adjudication as an epistemic accountability mechanism rather than a post-hoc validation step.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec27\" class=\"Section2\"\u003e \u003ch2\u003e5.3 RQ3: Diagnostic Utility of the E0\u0026ndash;E10 Taxonomy\u003c/h2\u003e \u003cp\u003eThe third research question assessed whether the E0\u0026ndash;E10 taxonomy generates meaningful distribution patterns and reflects calibrated epistemic accountability. The concentration of scores in the E0\u0026ndash;E4 range (70.2% of annotations) and the low frequency of E7\u0026ndash;E8 assignments (3.2%) indicate that the taxonomy effectively captures gradations of grounding quality. The absence of E5 scores suggests a relatively strict threshold between sufficient and insufficient grounding in structured cultural narratives; deviations either remain within safe limits or fall directly into severe categories.\u003c/p\u003e \u003cp\u003eNotably, E9 (abstention mechanism) accounted for 14.2% of annotations, appearing in both accepted and rejected contexts. Accepted E9 instances corresponded to appropriate epistemic restraint for out-of-context queries, demonstrating the system's capacity for boundary awareness (the ability to recognize knowledge limits and refrain from generating unsupported content) as conceptualized by Kapusta (\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e2025\u003c/span\u003e). Rejected E9 instances, by contrast, reflected procedural insufficiencies or non-substantive non-responses. This distinction confirms that abstention functions as a calibrated epistemic accountability indicator rather than a performance failure.\u003c/p\u003e \u003cp\u003eThe taxonomy thus serves dual purposes: as a diagnostic tool for identifying error types and as a calibration instrument that distinguishes between generative inability and intentional restraint. By enabling evaluators to verify evidence absence through artifact-level traceability, the framework renders abstention decisions accountable and transparent.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec28\" class=\"Section2\"\u003e \u003ch2\u003e5.4 RQ4: Contribution of Chunk Injection to Narrative Boundary Stability\u003c/h2\u003e \u003cp\u003eThe fourth research question explored the role of chunk injection in maintaining narrative boundary stability during iterative generation. The repeated testing of 14 question pairs revealed that 71.4% remained within the same severity range, with no instances escalating to attribution erosion (E7) or contradiction (E8). This stability indicates that explicit chunk injection into the model context acts as an epistemic constraint, limiting generative space and preventing extrapolation beyond retrieved evidence.\u003c/p\u003e \u003cp\u003eThese findings support Hickey's (2026) concept of \u003cem\u003econstitutive knowledge sources\u003c/em\u003e: epistemic trust is fostered not through algorithmic transparency but through authoritative, verifiable knowledge sources. Injected chunks serve precisely this function, transforming retrieval from probabilistic similarity matching to deterministic evidence provision. The observed stability also aligns with Kapusta's (2025) vision of two-level transparency, where high-level reasoning patterns and detailed factor explanations remain traceable. Annotators could consistently verify outputs across generations because chunk-level traceability provided sufficient evidence for in-depth validation.\u003c/p\u003e \u003cp\u003eThe Q13\u0026ndash;Q28 pair illustrates the stochastic nature of large language models while also demonstrating the effectiveness of epistemic constraints. Although the initial generation produced divergent scores (E6, E9, E10) leading to rejection, the repeated generation appropriately abstained (E9) and was accepted. Crucially, this variation did not escalate into attribution erosion or contradiction, confirming that even when outputs differ, they remain bounded by the retrieved evidence.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec29\" class=\"Section2\"\u003e \u003ch2\u003e5.5 RQ5: Traceability, Auditability, and Governance Alignment\u003c/h2\u003e \u003cp\u003eThe fifth research question evaluated how artifact preservation operationalizes transparency and accountability. The complete evaluation archive (Supplementary Material A) enables full reconstruction of every decision, from source chunks to consensus metadata. This exemplifies the OECD AI Principles' (2024) traceability-as-accountability instrument, allowing outcome analysis and inquiry resolution.\u003c/p\u003e \u003cp\u003eIn culturally sensitive domains where misrepresentation can harm heritage-owning communities, auditability becomes essential. Our framework operationalizes Shneiderman's (2022) Human-Centered AI vision by embedding human control and transparent reporting into the evaluation design. Annotators function not merely as validators but as holders of epistemic authority whose decisions are documented and accountable. This aligns with principles of moral authorship (Cristofaro \u0026amp; Ba\u0026ntilde;\u0026oacute;n Gomis, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e2026\u003c/span\u003e), where final decisions remain under human purview to preserve autonomy and responsibility.\u003c/p\u003e \u003cp\u003eBy providing an open dataset and comprehensive documentation, this research contributes to open science initiatives and enables replication, verification, and extension by the research community. The framework demonstrates how abstract governance principles (transparency, accountability, and reproducibility) can be translated into concrete, verifiable evaluation protocols.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec30\" class=\"Section2\"\u003e \u003ch2\u003e5.6 Limitations and Future Research\u003c/h2\u003e \u003cp\u003eSeveral limitations should be acknowledged. First, the study employed a single cultural corpus (the \u003cem\u003eAnde-Ande Lumut\u003c/em\u003e narrative). While this enabled controlled experimentation, generalization to other cultural contexts requires cross-corpus validation with diverse folklore, traditions, and narrative structures. Future research should apply the framework to multiple cultural datasets to assess its adaptability and identify culture-specific error patterns.\u003c/p\u003e \u003cp\u003eSecond, the evaluator panel consisted of three domain experts. Although inter-rater consensus was achieved through calibration and documented consensus adjudication, larger panels would enhance statistical robustness and allow for more nuanced analysis of evaluator effects. Future studies could incorporate evaluators from different cultural backgrounds to examine how interpretive frameworks influence attribution erosion judgments.\u003c/p\u003e \u003cp\u003eThird, while abstention (E9) is conceptualized as a positive indicator of boundary awareness, practical deployment may require balancing epistemic caution with user expectations. Users interacting with tourism chatbots typically anticipate informative responses; excessive abstention could degrade user experience. Future research should explore optimal trade-offs between epistemic restraint and user satisfaction, potentially developing adaptive abstention policies.\u003c/p\u003e \u003cp\u003eFourth, the current taxonomy, while granular, may benefit from refinement for partial automation. Developing concise, computationally detectable proxies for certain error categories could streamline evaluation without sacrificing diagnostic precision. Additionally, integrating alternative traceability methods such as GraphRAG (Edge et al., 2024) could address complex multi-hop questions requiring reasoning across distributed knowledge sources.\u003c/p\u003e \u003cp\u003eFinally, longitudinal studies are needed to assess how narrative boundary stability evolves over multiple interaction turns and whether repeated exposure to similar queries leads to attribution erosion accumulation. Such research would inform the design of self-correcting generative systems capable of maintaining cultural faithfulness over extended deployments.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec31\" class=\"Section2\"\u003e \u003ch2\u003e5.7 Summary\u003c/h2\u003e \u003cp\u003eCollectively, these findings demonstrate that Cultural Faithfulness can be systematically operationalized through the integration of artifact-level traceability, severity-differentiating diagnostics, structured human adjudication, and reproducible artifact preservation. The proposed framework advances evaluation practices beyond surface-level similarity metrics toward a unified model of epistemic accountability for generative systems in tourism and other culturally sensitive domains. By bridging smart tourism research, retrieval-augmented language modeling, structured human evaluation, and operational AI governance, this study provides a replicable template for assessing and ensuring cultural faithfulness in AI-mediated communication.\u003c/p\u003e \u003c/div\u003e"},{"header":"6. Conclusion","content":"\u003cp\u003eThis study introduces and empirically validates a traceable structured human adjudication framework for assessing Cultural Faithfulness in epistemically constrained Retrieval-Augmented Generation systems for tourism chatbots. Building upon UNESCO's (2003) normative framework recognizing oral traditions as intangible cultural heritage and OECD's (2024) governance principles positioning artifact-level traceability as an accountability instrument, this research addresses an urgent need for evaluation methodologies beyond conventional automated metrics.\u003c/p\u003e \u003cp\u003eEmpirical findings from 61 structured questions and 183 independent artifact-level traceability annotations reveal that 73.8% of responses were accepted after documented consensus adjudication, while 26.2% were rejected due to attribution erosion or severe grounding violations. Crucially, repeated testing showed no escalation into attribution erosion (E7) or contradiction (E8) categories, indicating the success of the boundary awareness mechanism embedded in the framework through chunk injection as contextual anchoring mechanisms.\u003c/p\u003e \u003cp\u003eFundamental Contributions\u003c/p\u003e \u003cp\u003eThis study presents four fundamental contributions that collectively shift the paradigm of AI evaluation from a technical-binary approach to epistemic accountability.\u003c/p\u003e \u003cp\u003eFirst, the study articulates the role of automated metrics as an analytical signal rather than a final determinant. Unlike conventional practices that treat ROUGE, BERTScore, or cosine similarity as proxies for truth, our framework positions these metrics as generators of initial diagnostic artifacts that enrich information for structured human adjudication. This design acknowledges the limitations of automated metrics in capturing nuanced fidelity aspects (Maynez et al., \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e2020\u003c/span\u003e) while empowering practical decision-makers with comprehensive information to fulfill their ethical and epistemic roles (Tschiatschek et al., \u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). Automated similarity filtering thus functions as a complementary layer within artifact-level traceability systems, not as a grounding validator.\u003c/p\u003e \u003cp\u003eSecond, the study redefines abstention and rejection as indicators of epistemic accountability. Within our framework, decisions to reject responses or invoke the abstention mechanism are not interpreted as system failures but as positive evidence of boundary awareness (the ability to recognize knowledge limits). This aligns with the vision of symbiotic epistemology (Kapusta, \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e2025\u003c/span\u003e), where reliable epistemic agents are those capable of acknowledging their ignorance, for instance by responding with \"\u003cem\u003eMaaf, saya tidak tahu\u003c/em\u003e\" (the local expression for \"I do not know\"), rather than providing answers with false confidence. In culturally codified contexts, abstention becomes an ethical act that prevents potential harm from cultural misrepresentation (Floridi et al., \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2018\u003c/span\u003e). The deliberate omission mechanism thus operationalizes epistemic accountability as a measurable construct.\u003c/p\u003e \u003cp\u003eThird, the study operationalizes reproducible artifact preservation across the entire pipeline as a foundation for accountability. The framework systematically produces and preserves artifacts at every stage: from injected source chunks, through severity-differentiating diagnostics (E0 to E10), attribution erosion indicators, to final documented consensus adjudication decisions. This artifact collection enables artifact-level traceability, auditability, and epistemic accountability of every decision. It represents a concrete implementation of transparency as an accountability instrument (OECD, \u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e2024\u003c/span\u003e) and constitutive knowledge as the foundation for epistemic trust in opaque systems (Hickey, \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2026\u003c/span\u003e). The reproducible artifact preservation framework ensures that every evaluation decision is traceable to specific textual evidence through artifact-level traceability systems.\u003c/p\u003e \u003cp\u003eFourth, the study introduces a granular error taxonomy from E0 to E10 as a collaborative reasoning instrument for structured human adjudication. This severity-differentiating taxonomy elevates evaluation from binary judgments of correctness to rich epistemic diagnostics of narrative boundary stability. It enables annotators to function as co-reasoners (Salloch \u0026amp; Eriksen, \u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e2024\u003c/span\u003e) who collaboratively interpret evidence, categorize attribution erosion typologies, and achieve legitimate documented consensus. This process generates moral authorship (Cristofaro \u0026amp; Ba\u0026ntilde;\u0026oacute;n Gomis, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e2026\u003c/span\u003e), where final decisions remain genuinely human, not merely formalities within algorithmic loops. The taxonomy thus serves as both diagnostic instrument and governance mechanism for culturally rooted narrative agents.\u003c/p\u003e \u003cp\u003eImplications and Future Research\u003c/p\u003e \u003cp\u003eThis study not only proposes an alternative evaluation method but establishes a novel epistemic accountability paradigm for generative AI. The paradigm emphasizes that in culturally sensitive domains like tourism, Cultural Faithfulness cannot be reduced to similarity scores or statistical thresholds. It necessitates active human engagement as epistemic authorities, supported by artifact-level traceability systems and structured collaborative reasoning processes.\u003c/p\u003e \u003cp\u003eThe finding that 26.2% of responses were rejected demonstrates that Cultural Faithfulness constitutes a dimension far more demanding than mere factual accuracy. This figure does not signify system weakness but rather reflects the complexity of culturally codified heritage resisting simplification into vector representations. Crucially, it reminds us that AI development for cultural tourism must collaborate with cultural custodians, not merely rely on technical optimization. The rejection rate operationalizes the gap between technical performance and epistemic accountability.\u003c/p\u003e \u003cp\u003eFuture research directions include:\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eTaxonomy transferability across diverse cultural domains to validate the severity-differentiating diagnostics framework.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eOptimizing the balance between automated similarity filtering and structured human adjudication to enhance efficiency without compromising epistemic accountability.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eIntegrating with alternative traceability architectures such as GraphRAG (Edge et al., 2024) to address complex multi-hop queries while maintaining narrative boundary stability.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eDeveloping governance frameworks that embed reproducible artifact preservation into institutional practices for cultural heritage protection.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cp\u003eMost importantly, this framework offers a governance model where technology does not replace but rather strengthens human roles as cultural heritage stewards in the digital era. By operationalizing transparency and accountability principles through artifact-level traceability systems, structured human adjudication paradigms, and reproducible artifact preservation mechanisms, this research contributes to a broader vision\u0026mdash;developing generative systems that are not only technically intelligent but also culturally faithful and epistemically accountable.\u003c/p\u003e "},{"header":"Declarations","content":"\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eAll authors contributed to the study conception and design. Hindun Nurhidayati was responsible for the cultural tourism conceptual framework and dataset preparation, drawing upon the foundational book Dilema penutur dalam pariwisata budaya co-authored by all contributors. The communication and language aspects were overseen by Zulin Nurchayati. Sugiarto, as an independent researcher, led the technical implementation of AI models, retrieval architecture design, and data management. All authors read and approved the final manuscript.\u003c/p\u003e\u003ch2\u003eAcknowledgement\u003c/h2\u003e\u003cp\u003eThe authors express their gratitude to colleagues for their valuable discussions and support throughout this research. Special appreciation is also extended to the evaluators who participated in the annotation and adjudication processes.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe dataset generated and analyzed during the current study is owned by the authors and has been deposited in the Zenodo repository under a temporary embargo to ensure double blind review. The dataset will be made publicly available upon publication at: https://doi.org/10.5281/zenodo.18603955During the peer review process, the dataset is provided as Supplementary Material A (file name: Supplementary_Material_A_Cultural_Faithfulness_Dataset.pdf) under a generic title that does not reveal author identities. This material contains all evaluation artifacts necessary for replicating the study.Note to the EditorThe authors wish to inform the editor that the two references cited in the manuscript with the author names Nurhidayati, Nurchayati, \u0026amp; Sugiarto (2025, 2026) are works authored by the same research team. These publications, comprising a book and a conceptual preprint, represent a deliberate research trajectory that establishes the theoretical and empirical foundation for the present study. This disclosure is made to avoid any perception of self‑plagiarism and to ensure full transparency.\u003c/p\u003e\n\u003ch3\u003e\u003cstrong\u003eFunding Declarations\u003c/strong\u003e\u003c/h3\u003e\n\u003cp\u003eThe authors declare that no funds, grants, or other support were received during the preparation of this manuscript.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eAbid A, Farooqi M, Zou J (2021) Persistent anti-Muslim bias in large language models. In: \u003cem\u003eProceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society\u003c/em\u003e, pp. 298\u0026ndash;306. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1145/3461702.3462624\u003c/span\u003e\u003cspan address=\"10.1145/3461702.3462624\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAmershi S, Weld D, Vorvoreanu M, Fourney A, Nushi B, Collisson P, Suh J, Iqbal S, Bennett PN, Inkpen K, Teevan J, Kikin-Gil R, Horvitz E (2019) Guidelines for human-AI interaction. In: \u003cem\u003eProceedings of the 2019 CHI Conference on Human Factors in Computing Systems\u003c/em\u003e, Paper No. 3, pp. 1\u0026ndash;13. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1145/3290605.3300233\u003c/span\u003e\u003cspan address=\"10.1145/3290605.3300233\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eArtstein R, Poesio M (2008) Inter-coder agreement for computational linguistics. Comput Linguistics 4555\u0026ndash;596. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1162/coli.07-034-R2\u003c/span\u003e\u003cspan address=\"10.1162/coli.07-034-R2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAsai A, Salehi M, Peters ME, Hajishirzi H (2022) ATTEMPT: Parameter-efficient multi-task tuning via attentional mixtures of soft prompts. \u003cem\u003eProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing\u003c/em\u003e, 6655\u0026ndash;6672. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.18653/v1/2022.emnlp-main.446\u003c/span\u003e\u003cspan address=\"10.18653/v1/2022.emnlp-main.446\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBender EM, Gebru T, McMillan-Major A, Shmitchell S (2021) On the dangers of stochastic parrots: Can language models be too big? \u003cem\u003eProceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency\u003c/em\u003e, 610\u0026ndash;623. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1145/3442188.3445922\u003c/span\u003e\u003cspan address=\"10.1145/3442188.3445922\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBirhane A, Kalluri P, Card D, Agnew W, Dotan R, Bao M (2022) The values encoded in machine learning research. In: \u003cem\u003eProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (FAccT '22)\u003c/em\u003e, pp. 173\u0026ndash;184. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1145/3531146.3533083\u003c/span\u003e\u003cspan address=\"10.1145/3531146.3533083\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBlodgett SL, Barocas S, Daum\u0026eacute; H III, Wallach H (2020) Language (technology) is power: A critical survey of bias in NLP. \u003cem\u003eProceedings of the 58th Annual Meeting of the Association for Computational Linguistics\u003c/em\u003e, 5454\u0026ndash;5476. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.18653/v1/2020.acl-main.485\u003c/span\u003e\u003cspan address=\"10.18653/v1/2020.acl-main.485\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBorgeaud S, Mensch A, Hoffmann J, Cai T, Rutherford E, Casas K, Sifre L (2022) Improving language models by retrieving from trillions of tokens. \u003cem\u003eProceedings of the 39th International Conference on Machine Learning\u003c/em\u003e, PMLR 162, 2206\u0026ndash;2240. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://proceedings.mlr.press\u003c/span\u003e\u003cspan address=\"https://proceedings.mlr.press\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCampbell JL, Quincy C, Osserman J, Pedersen OK (2013) Coding in-depth semistructured interviews: Problems of unitization and intercoder reliability and agreement. Sociol Methods Res 3294\u0026ndash;320. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1177/0049124113500475\u003c/span\u003e\u003cspan address=\"10.1177/0049124113500475\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCao Y, Zhou L, Lee S, Cabello L, Chen M, Hershcovich D (2023) Assessing cross-cultural alignment between ChatGPT and human societies: An empirical study. *Proceedings of the First Workshop on Cross-Cultural Considerations in NLP (C3NLP)*. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://aclanthology.org/2023.c3nlp-1.3\u003c/span\u003e\u003cspan address=\"https://aclanthology.org/2023.c3nlp-1.3\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCaric H, Mandić A, Sever I (2026) A six-phase AI-expert framework for evaluating policy coherence in sustainable tourism. J Inform Technol Tourism 6. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s40558-025-00352-0\u003c/span\u003e\u003cspan address=\"10.1007/s40558-025-00352-0\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. *28*\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCohen J (1960) A coefficient of agreement for nominal scales. Educ Psychol Meas 20(1):37\u0026ndash;46. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1177/001316446002000104\u003c/span\u003e\u003cspan address=\"10.1177/001316446002000104\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCristofaro M, Ba\u0026ntilde;\u0026oacute;n Gomis G (2026) Dancing with the algorithm: A framework to navigate knowledge and autonomy in AI-assisted managerial decisions. J Knowl Manage. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1108/JKM-06-2025-0870\u003c/span\u003e\u003cspan address=\"10.1108/JKM-06-2025-0870\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCruz M, Jardim B, de Castro Neto M (2025) Lisa: A touristic chatbot for Lisbon. J Inform Technol Tourism 1153\u0026ndash;1183. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s40558-025-00339-x\u003c/span\u003e\u003cspan address=\"10.1007/s40558-025-00339-x\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. *27*\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEdge D, Trinh H, Cheng N, Bradley J, Chao A, Mody A, Truitt S, Metropolitansky D, Ness RO, Larson J (2025) \u003cem\u003eFrom local to global: A graph RAG approach to query-focused summarization\u003c/em\u003e (Version 2). arXiv. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.48550/arXiv.2404.16130\u003c/span\u003e\u003cspan address=\"10.48550/arXiv.2404.16130\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFleiss JL (1971) Measuring nominal scale agreement among many raters. \u003cem\u003ePsychological Bulletin\u003c/em\u003e, *76*(5), 378\u0026ndash;382. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1037/h0031619\u003c/span\u003e\u003cspan address=\"10.1037/h0031619\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFloridi L, Cowls J, Beltrametti M et al (2018) AI4People\u0026mdash;An ethical framework for a good AI society: Opportunities, risks, principles, and recommendations. \u003cem\u003eMinds and Machines\u003c/em\u003e, *28*, 689\u0026ndash;707. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s11023-018-9482-5\u003c/span\u003e\u003cspan address=\"10.1007/s11023-018-9482-5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGao Y, Xiong Y, Gao X, Jia H, Pan Y, Bi Y, Dai Y, Sun J, Wang H, Wang H (2024) Retrieval-augmented generation for large language models: A survey. arXiv preprint. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://arxiv.org/abs/2312.10997\u003c/span\u003e\u003cspan address=\"https://arxiv.org/abs/2312.10997\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGretzel U (2011) Intelligent systems in tourism: A social science perspective. Annals Tourism Res 38(3):757\u0026ndash;781. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.annals.2011.04.014\u003c/span\u003e\u003cspan address=\"10.1016/j.annals.2011.04.014\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGretzel U, Sigala M, Xiang Z, Koo C (2015) Smart tourism: Foundations and developments. Electron Markets 25(3):179\u0026ndash;188. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s12525-015-0196-8\u003c/span\u003e\u003cspan address=\"10.1007/s12525-015-0196-8\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuu K, Lee K, Tung Z, Pasupat P, Chang MW (2020) Retrieval augmented language model training. In Proceedings of the 37th International Conference on Machine Learning (Vol. 119, pp. 3929\u0026ndash;3938). PMLR. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://proceedings.mlr.press\u003c/span\u003e\u003cspan address=\"https://proceedings.mlr.press\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHershcovich D, Frank S, Lent H, de Lhoneux M, Abdou M, Brandl S, Bugliarello E, Piqueras C, Chalkidis L, Cui I, Fierro R, Margatina C, Rust K, P., S\u0026oslash;gaard A (2022) Challenges and strategies in cross-cultural NLP. \u003cem\u003eProceedings of the 60th Annual Meeting of the Association for Computational Linguistics\u003c/em\u003e, 6997\u0026ndash;7013. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.18653/v1/2022.acl-long.482\u003c/span\u003e\u003cspan address=\"10.18653/v1/2022.acl-long.482\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHickey C (2026) Constitutive knowledge sources: An institutional approach to epistemic trust in opaque AI systems. \u003cem\u003eAI and Ethics\u003c/em\u003e, *6*, 58. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s43681-025-00930-2\u003c/span\u003e\u003cspan address=\"10.1007/s43681-025-00930-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHofstede G, Hofstede GJ, Minkov M (2010) Cultures and organizations: Software of the mind: Intercultural cooperation and its importance for survival, 3rd edn. McGraw-Hill\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJi Z, Han S, Yu S (2023) Survey of hallucination in natural language generation. ACM-CSUR 55(12):1\u0026ndash;38. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1145/3571730\u003c/span\u003e\u003cspan address=\"10.1145/3571730\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJobin A, Ienca M, Vayena E (2019) The global landscape of AI ethics guidelines. Nat Mach Intell 9389\u0026ndash;399. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s42256-019-0088-2\u003c/span\u003e\u003cspan address=\"10.1038/s42256-019-0088-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKapusta J (2025) SynLang and symbiotic epistemology: A manifesto for conscious human-AI collaboration. arXiv preprint. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://arxiv.org/abs/2507.21067\u003c/span\u003e\u003cspan address=\"https://arxiv.org/abs/2507.21067\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKim JH, Kim J, Park J, Kim C, Jhang J, King B (2025) When ChatGPT gives incorrect answers: The impact of inaccurate information by generative AI on tourism decision-making. J Travel Res 64(1):51\u0026ndash;73. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1177/00472875231212996\u003c/span\u003e\u003cspan address=\"10.1177/00472875231212996\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKrippendorff K (2019) \u003cem\u003eContent analysis: An introduction to its methodology\u003c/em\u003e (4th ed.). SAGE Publications. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.4135/9781071878781\u003c/span\u003e\u003cspan address=\"10.4135/9781071878781\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLandis JR, Koch GG (1977) The measurement of observer agreement for categorical data. Biometrics *33* 1159\u0026ndash;174. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.2307/2529310\u003c/span\u003e\u003cspan address=\"10.2307/2529310\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLaux J, Wachter S, Mittelstadt B (2024) Trustworthy artificial intelligence and the European Union AI act: On the conflation of trustworthiness and acceptability of risk. Regul Gov 18(1):3\u0026ndash;32. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1111/rego.12512\u003c/span\u003e\u003cspan address=\"10.1111/rego.12512\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLee N, Bang Y, Madotto A, Fung P (2025) Cultural alignment test (Hofstede's CAT): Revealing the struggle of Large Language Models in comprehending cultural values. On \u003cem\u003eProceedings of the 31st International Conference on Computational Linguistics (COLING 2025)\u003c/em\u003e. Association for Computational Linguistics. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://aclanthology.org/2025.coling-main.567/\u003c/span\u003e\u003cspan address=\"https://aclanthology.org/2025.coling-main.567/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLewis P, Perez E, Piktus A, Petroni F, Karpukhin V, Goyal N, K\u0026uuml;ttler H, Lewis M, Yih WT, Rockt\u0026auml;schel T, Riedel S, Kiela D (2020) Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459\u0026ndash;9474. proceedings.neurips.cc\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLin C-Y (2004) ROUGE: A package for automatic evaluation of summaries. \u003cem\u003eProceedings of the Association for Computational Linguistics Workshop\u003c/em\u003e, 74\u0026ndash;81. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://aclanthology.org/W04-1013/\u003c/span\u003e\u003cspan address=\"https://aclanthology.org/W04-1013/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu CC, Gurevych I, Korhonen A (2025) Culturally aware and adapted NLP: A taxonomy and a survey of the state of the art. Trans Association Comput Linguistics 13:652\u0026ndash;689. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1162/tacl_a_00760\u003c/span\u003e\u003cspan address=\"10.1162/tacl_a_00760\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLucy L, Bamman D (2021) Gender and representation bias in GPT-3 generated stories. In: \u003cem\u003eProceedings of the Third Workshop on Narrative Understanding\u003c/em\u003e, pp. 48\u0026ndash;55. Association for Computational Linguistics. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.18653/v1/2021.nuse-1.5\u003c/span\u003e\u003cspan address=\"10.18653/v1/2021.nuse-1.5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMaynez J, Narayan S, Bohnet B, McDonald R (2020) On faithfulness and factuality in abstractive summarization. \u003cem\u003eProceedings of the 58th Annual Meeting of the Association for Computational Linguistics\u003c/em\u003e, 1906\u0026ndash;1919. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://aclanthology.org/2020.acl-main.173\u003c/span\u003e\u003cspan address=\"https://aclanthology.org/2020.acl-main.173\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMittelstadt BD, Allo P, Taddeo M et al (2016) The ethics of algorithms. Big Data Soc. *3*(2 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1177/2053951716679679\u003c/span\u003e\u003cspan address=\"10.1177/2053951716679679\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMohamed S, Png MT, Isaac W (2020) Decolonial AI: Decolonial theory as sociotechnical foresight in artificial intelligence. Philos Technol 4659\u0026ndash;684. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s13347-020-00405-8\u003c/span\u003e\u003cspan address=\"10.1007/s13347-020-00405-8\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNurhidayati H, Nurchayati Z, Sugiarto (2025) Dilema penutur dalam pariwisata budaya. Deepublish\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNurhidayati H, Nurchayati Z, Sugiarto (2026) Lima teori konseptual untuk integrasi AI dan komunikasi budaya dalam pariwisata. \u003cem\u003ePreprint at Zenodo\u003c/em\u003e. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.5281/zenodo.18335863\u003c/span\u003e\u003cspan address=\"10.5281/zenodo.18335863\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOECD (2019) OECD principles on artificial intelligence. OECD Publishing. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://oecd.ai/en/ai-principles\u003c/span\u003e\u003cspan address=\"https://oecd.ai/en/ai-principles\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOECD (2024) \u003cem\u003eOECD Artificial Intelligence Principles 2024\u003c/em\u003e. OECD Publishing. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://oecd.ai/en/ai-principles\u003c/span\u003e\u003cspan address=\"https://oecd.ai/en/ai-principles\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRen J, Xu Y, Wang X, Li W, Wang A, Ma W, Liu Y (2025) Towards transparent RAG: Fostering evidence traceability in LLM generation via reinforcement learning. arXiv preprint. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.48550/arXiv.2505.13258\u003c/span\u003e\u003cspan address=\"10.48550/arXiv.2505.13258\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSalemi A, Zamani H (2024) Evaluating retrieval quality in retrieval-augmented generation. \u003cem\u003eProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval\u003c/em\u003e, 2395\u0026ndash;2400. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1145/3626772.3657957\u003c/span\u003e\u003cspan address=\"10.1145/3626772.3657957\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSalloch S, Eriksen A (2024) What are humans doing in the loop? Co-reasoning and practical judgment when using machine learning-driven decision aids. Am J Bioeth 24(9):67\u0026ndash;78. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1080/15265161.2024.2353800\u003c/span\u003e\u003cspan address=\"10.1080/15265161.2024.2353800\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShneiderman B (2022) Human-centered AI. Oxford University Press. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1093/oso/9780192845290.001.0001\u003c/span\u003e\u003cspan address=\"10.1093/oso/9780192845290.001.0001\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSigala M (2018) New technologies in tourism: From multi-disciplinary to anti-disciplinary advances and trajectories. Tourism Management Perspectives, *25*, 151\u0026ndash;155. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.tmp.2017.12.003\u003c/span\u003e\u003cspan address=\"10.1016/j.tmp.2017.12.003\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStilgoe J, Owen R, Macnaghten P (2013) Developing a framework for responsible innovation. \u003cem\u003eResearch Policy\u003c/em\u003e, *42*(9), 1568\u0026ndash;1580. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.respol.2013.05.008\u003c/span\u003e\u003cspan address=\"10.1016/j.respol.2013.05.008\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTafjord O, Dalvi B, Clark P (2021) ProofWriter: Generating implications, proofs, and abductive statements over natural language. \u003cem\u003eFindings of the Association for Computational Linguistics: ACL-IJCNLP 2021\u003c/em\u003e, 3621\u0026ndash;3634. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.18653/v1/2021.findings-acl.317\u003c/span\u003e\u003cspan address=\"10.18653/v1/2021.findings-acl.317\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTing-Toomey S (2015) Intercultural and intergroup communication competence: Toward an integrative perspective. In G. Rickheit \u0026amp; H. Strohner (Eds.), \u003cem\u003eHandbook of communication competence\u003c/em\u003e (pp. 503\u0026ndash;538). De Gruyter Mouton. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1515/9783110317459-021\u003c/span\u003e\u003cspan address=\"10.1515/9783110317459-021\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTsamados A, Aggarwal N, Cowls J, Morley J, Roberts H, Taddeo M, Floridi L (2022) The ethics of algorithms: key problems and solutions. AI \u0026amp; Society, *37*, pp 215\u0026ndash;230. 1\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s00146-021-01154-8\u003c/span\u003e\u003cspan address=\"10.1007/s00146-021-01154-8\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTschiatschek S, Stamboliev E, Schmude T, Coeckelbergh M, Koesten L (2024) Challenging the human-in-the-loop in algorithmic decision-making. arXiv preprint. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.48550/arXiv.2405.10706\u003c/span\u003e\u003cspan address=\"10.48550/arXiv.2405.10706\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eUNESCO (2003) Convention for the safeguarding of the intangible cultural heritage. UNESCO. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ich.unesco.org/en/convention\u003c/span\u003e\u003cspan address=\"https://ich.unesco.org/en/convention\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eUNESCO (2025) AI and the future of education: Disruptions, dilemmas and directions. UNESCO Publishing. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://unesdoc.unesco.org/ark:/48223/pf0000395236\u003c/span\u003e\u003cspan address=\"https://unesdoc.unesco.org/ark:/48223/pf0000395236\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang T, Kishore V, Wu F, Weinberger KQ, Artzi Y (2020) BERTScore: Evaluating text generation with BERT. In \u003cem\u003eProceedings of the 8th International Conference on Learning Representations\u003c/em\u003e (ICLR 2020). OpenReview.net. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://openreview.net/forum?id=SkeHuCVFDr\u003c/span\u003e\u003cspan address=\"https://openreview.net/forum?id=SkeHuCVFDr\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Cultural Faithfulness, Artifact-Level Traceability, Attribution Erosion, Structured Human Adjudication, Narrative Boundary Stability, Epistemic Accountability","lastPublishedDoi":"10.21203/rs.3.rs-8994226/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8994226/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe integration of large language model-based chatbots into tourism contexts has sparked critical discussions about Cultural Faithfulness, especially concerning the accurate representation of intangible heritage. While Retrieval-Augmented Generation (RAG) enhances factual accuracy, current evaluation methods predominantly depend on automated similarity metrics and seldom incorporate structured human adjudication, often conflating semantic coherence with epistemic validity.\u003c/p\u003e \u003cp\u003eTo address these limitations, this study proposes a Human-in-the-Loop evaluation framework for traceable RAG systems in tourism chatbots. The framework combines chunk-level retrieval traceability, a granular narrative error taxonomy (E0\u0026ndash;E10) designed to capture varying degrees of attribution erosion, automated similarity filtering, and human-based documented consensus adjudication into a cohesive protocol. By treating retrieval as an epistemic constraint on generative processes and operationalizing abstention as a measure of boundary awareness, the framework establishes rigorous evaluation criteria.\u003c/p\u003e \u003cp\u003eEmpirical validation was conducted using 61 structured questions derived from a corpus of Indonesian cultural narratives, generating 183 independent annotations. Analysis revealed that 73.8% of responses met acceptance criteria after artifact-level review, while 26.2% were excluded due to grounding violations. Notably, high-severity errors were confined to rejected cases, and iterative testing confirmed no progression into hallucination or contradiction categories.\u003c/p\u003e \u003cp\u003eThese results confirm that Cultural Faithfulness can be systematically achieved through traceable retrieval mechanisms, structured human validation, and governance-aligned artifact preservation. This research extends evaluation methodologies beyond superficial similarity metrics, advancing a unified model of epistemic accountability for generative systems in tourism applications.\u003c/p\u003e","manuscriptTitle":"Cultural Faithfulness in Tourism Chatbots: A Structured Human Adjudication Framework for Traceable Retrieval-Augmented Generation","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-18 10:09:44","doi":"10.21203/rs.3.rs-8994226/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"b392a459-583b-4c8e-b180-a7f484651c03","owner":[],"postedDate":"March 18th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-04-14T16:55:24+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-18 10:09:44","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8994226","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8994226","identity":"rs-8994226","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.