Research on a Privacy-Aware Data Processing Framework for Large-Scale Distributed Systems | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Research on a Privacy-Aware Data Processing Framework for Large-Scale Distributed Systems Jingzhou Tao, Zijing Liu, Ruoxi Lyu This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8923292/v1 This work is licensed under a CC BY 4.0 License Status: Under Revision Version 1 posted 11 You are reading this latest preprint version Abstract The widespread adoption of large-scale distributed systems has promoted the efficient circulation and processing of multi-source data; however, the risk of data privacy leakage has become a core bottleneck constraining their development. This paper focuses on theoretical research on privacy-aware data processing. Based on the data processing characteristics of distributed systems and the core requirements of privacy protection, it constructs a privacy-aware data processing framework at the theoretical level, clarifies the framework’s core layers and theoretical connotations, and systematically analyzes the theoretical linkages among key enabling technologies. Through a literature review of existing related work, this paper identifies the current research status and gaps in distributed data processing and privacy protection, providing theoretical support for the collaborative optimization of data processing and privacy protection in large-scale distributed systems. Physical sciences/Engineering Physical sciences/Mathematics and computing large-scale distributed systems privacy awareness data processing Figures Figure 1 1 Introduction 1.1 Research Background and Significance With the development of technologies such as cloud computing and the Internet of Things, large-scale distributed systems have become the core of data processing across multiple domains. However, due to characteristics such as heterogeneous nodes and decentralization, the entire data lifecycle faces severe privacy leakage risks; more than 60% of security incidents are related to this, constraining the industry’s development. Most existing research focuses on system efficiency and resource scheduling, while theoretical research on integrating privacy protection with data processing remains insufficient and lacks a complete system. This study establishes a privacy-aware processing framework to fill the theoretical gap. It can provide design references for key fields such as energy and healthcare, ensure data security and compliant use, promote standardized industrial development, and enhance China’s technological competitiveness in related areas, thereby offering important theoretical and practical significance. 1.2 Research Status at Home and Abroad 1.2.1 Research Status Abroad Research abroad on distributed data processing and privacy protection started relatively early, mainly from three perspectives: heterogeneous data processing, integration of privacy computing, and multi-node collaborative optimization, achieving notable results. In data processing, Huang Y et al. [2] proposed a physics-informed pseudo-labeling unsupervised learning theory that combines physical models with unsupervised learning, effectively handling heterogeneous data and reducing exposure of sensitive data through pseudo-labels. Zhang R et al. [3] created a distributed peer-to-peer transaction framework for urban energy Internet scenarios, reducing the risk of centralized exposure of transaction data via preference awareness and distributed negotiation. Liu Y et al. [4] proposed distributed online fusion learning, avoiding centralized transfer of sensitive data through collaborative multi-node training. In addition, parallel processing models based on data sharding and load balancing provide support for large-scale data processing [10]. In privacy protection, the core lies in integrating cryptographic technologies with systems to form theories such as homomorphic encryption and secure multi-party computation, enabling direct computation on encrypted data and joint computation without raw data. Lightweight algorithms are suitable for resource-constrained nodes [2][3], and adaptive theories adjust strategies according to node states and data sensitivity [4]. Limitations include: privacy protection theories are highly targeted and lack a universal framework; research on privacy leakage propagation mechanisms is scarce; and theories for collaboratively optimizing privacy and efficiency under large-scale heterogeneous nodes remain incomplete. 1.2.2 Research Status in China Domestic research focuses on data processing, privacy protection, and resource scheduling, laying a foundation for framework construction. In data processing, Que Hang et al. [1] proposed a resource-virtualization architecture to guide ISAC multi-source data processing; Wan Kaicheng [10] improved processing efficiency through data-locality scheduling and storage optimization; Yang Xiaolan [13] established a layered architecture for massive parallel data processing; Shu Xinyue [12] addressed data-center scheduling through load balancing and energy-consumption optimization; and intelligent optimization breakthroughs have also appeared in distributed inference [5][14]. In privacy protection, He Lili et al. [7] created a secure framework for mobile crowdsensing; Kou Lan et al. [15] combined compressive sensing with hash functions to balance privacy and accuracy; Tang Ronghua [11] and Ni Huiming [6] proposed methods for encrypted transmission in campus networks and secure sharing of medical data, respectively. Collaborative optimization work also includes Wei Hao et al.’s optimization of integrated sensing-and-communication network scheduling and Zhu Meiyi’s model compression to reduce sensitive data transmission. Limitations include: insufficiently deep integration between privacy protection and data processing, lacking a full-lifecycle awareness system; weak top-level design; and a lack of systematic analysis of collaborative optimization between privacy and performance. 1.3 Research Methods This study centers on theoretical construction and logical analysis, adopting multiple research methods to ensure scientific rigor and systematicity: Literature review: organizing research on large-scale distributed systems, data processing theories, and privacy protection technologies at home and abroad, with a focus on analyzing the 15 provided references, summarizing progress and gaps to support the establishment of the theoretical framework. Logical analysis: following the logical chain of “problem–cause–measure” to analyze manifestations and mechanisms of privacy leakage in distributed data processing, thereby deriving the construction approach and theoretical system of the privacy-aware data processing framework. Theoretical modeling: leveraging multidisciplinary foundations such as distributed systems theory, privacy computing theory, and data processing theory to establish a theoretical model, determining the layered structure, main functions, and collaboration modes of the framework. Comparative analysis: comparing theoretical characteristics of different distributed data processing models and privacy protection technologies to identify applicable scenarios, strengths, and weaknesses, providing theoretical basis for selecting and integrating key technologies within the framework. 2 Fundamental Theories of Data Processing and Privacy Protection in Large-Scale Distributed Systems 2.1 Fundamental Theories of Large-Scale Distributed Systems 2.1.1 Definition and Characteristics of Distributed Systems A distributed system consists of multiple independent nodes geographically dispersed and connected via a network. Nodes collaborate through coordinated work to achieve common objectives. Its core characteristics include decentralization (no unified control center), peer-to-peer communication among nodes, heterogeneity (differences in hardware, operating systems, and software platforms), dynamicity (nodes can join or leave at any time and the network topology changes dynamically), reliability (fault tolerance improved through multi-node redundancy), and scalability (processing and storage capacity can be enhanced by adding nodes). Large-scale distributed systems refer to systems with thousands or even tens of thousands of nodes and PB-level or larger data volumes; scenarios are more complex and impose higher requirements on coordination efficiency and reliability. 2.1.2 Distributed System Architecture Theory Distributed system architecture theory lays the foundation for data processing and privacy protection. Mainstream architectures include layered architectures, microservices architectures, and peer-to-peer network architectures. Their theoretical characteristics and application scenarios differ significantly, as shown in Table 2 − 1. Table 2 − 1 Comparison of Mainstream Distributed System Architecture Theories Architecture Type Core Idea Key Components Advantages Disadvantages Application Scenario Layered architecture Layered collaboration of functions; standardized interfaces Data/Network/Processing/Application layers Highly modular; easy to maintain and extend; privacy easier to protect High inter-layer communication overhead; high latency Large-scale data storage and batch processing Microservices architecture Split into independent services; lightweight collaboration Service registry/config center; API gateway, etc. Flexible and scalable; limited blast radius of failures Complex service dependencies; difficult coordination Heterogeneous data processing Peer-to-peer network architecture Equal nodes; distributed collaboration Peer nodes; distributed hash tables, etc. Decentralized; strong fault tolerance; low latency Difficult node management; no unified privacy control Real-time data sharing Edge–cloud collaborative architecture Real-time edge processing; cloud-side optimization Edge layer; network layer; cloud layer Strong real-time performance; less transmission; lower privacy risk Limited resources at edge nodes Low-latency scenarios Architecture design for large-scale distributed systems must balance data processing efficiency, privacy protection requirements, and resource constraints. In practice, a layered edge–cloud collaborative architecture is often adopted to achieve coordination between local real-time processing and global optimization. 2.1.3 Distributed Data Processing Model Theory Distributed data processing models are the core for efficient processing of massive data. Depending on real-time requirements and data characteristics, mainstream models include batch processing, stream processing, and hybrid processing models (Lambda architecture and Kappa architecture). Their theoretical characteristics and applicable scenarios are as follows: The core theory of batch processing is to perform offline batch processing on massive static data, improving efficiency through data sharding and parallel computing. A typical representative is the MapReduce model, whose theoretical process includes the Map phase (data partitioning and local computation), Shuffle phase (data distribution and sorting), and Reduce phase (global aggregation computation). Advantages include high throughput and high accuracy, suitable for non-real-time scenarios such as big-data analytics and report generation. The disadvantage is high latency, which cannot meet real-time processing requirements. Stream processing: the theoretical core is real-time processing of continuously generated dynamic data streams. Data is processed as it arrives without being stored before computation. Typical representatives include Spark Streaming and Flink, characterized by low latency (millisecond level), event-driven processing, and incremental computation. It is suitable for scenarios requiring real-time processing such as network situational awareness and real-time monitoring, but may suffer from accuracy constraints due to speed requirements and high resource consumption. Hybrid processing models: the Lambda architecture combines the advantages of batch and stream processing—the batch layer processes full data to ensure accuracy, the stream layer processes real-time data to ensure low latency, and the serving layer provides unified query services. The Kappa architecture simplifies Lambda by using a single stream-processing layer to process full data and using data replay to ensure accuracy. Hybrid models are suitable for scenarios requiring both real-time performance and accuracy, such as distributed energy dispatch and urban traffic monitoring. The theoretical performance comparison of these models is shown in Table 2 – 2 . Table 2 2 Theoretical Performance Comparison of Distributed Data Processing Models Performance Metric Batch Processing Stream Processing Lambda Architecture Kappa Architecture Theoretical processing latency High Low Medium Medium–low Theoretical throughput High Medium High Medium–high Processing accuracy High Medium High Medium–high Theoretical resource consumption Medium High High Medium Privacy-protection friendliness Medium High Medium–high Medium–high 2.2 Core Theories of Privacy Protection Privacy refers to the data subject’s control over sensitive information—i.e., the right for information not to be illegally collected, used, or disclosed. In large-scale distributed systems, privacy information includes users’ personal information, business-sensitive data, and system operational data, among others. Privacy leakage refers to privacy information being obtained, used, or disclosed without authorization, including direct leakage via theft of raw data and indirect leakage inferred through correlation analysis. Cryptography is the theoretical foundation of privacy protection, providing the mathematical principles for encryption protection in distributed systems. It mainly includes three theories: symmetric encryption (e.g., AES, RC4), where encryption and decryption share a key—high efficiency and low resource consumption, suitable for large-scale data transmission and storage, but key distribution and management are difficult in multi-node distributed environments and leakage risk is high; asymmetric encryption (e.g., RSA, ECC), which uses public-key encryption and private-key decryption—simpler key management without a secure channel, but low efficiency and unsuitable for real-time processing of massive data; homomorphic encryption supports computation directly on encrypted data, enabling “data usable but not visible,” serving as a core enabler for privacy computing, but with high computational complexity and resource consumption and is currently difficult to adapt to large-scale real-time scenarios. Privacy computing theory concerns balancing privacy protection and value extraction. Key technologies include: federated learning, which enables multi-node joint training by sharing model parameters and includes horizontal, vertical, and transfer variants—no raw data transmission and suitable for modeling with large numbers of nodes, but with high communication overhead and slow convergence; secure multi-party computation (SMPC), which uses secret sharing and related techniques to enable joint computation without data disclosure—suitable for small-scale collaborative decision-making such as financial reconciliation, but with high computational complexity and poor real-time performance; differential privacy, which protects individual privacy by adding noise and includes centralized and local variants—simple and low resource consumption, suitable for ultra-large-scale data collection, but noise affects data usability. A comparison of theoretical characteristics of privacy computing technologies is shown in (Table 2 –3 Detailed Comparison of Theoretical Characteristics of Privacy Computing Technologies Technology Type Privacy Protection Strength Computational Efficiency Communication Overhead Data Usability Applicable Node Scale Typical Application Scenario Federated learning High Medium Medium–high High Large-scale Distributed data modeling, e.g., collaborative medical data analysis Secure multi-party computation Very high Low High Very high Small-scale Distributed collaborative decision-making, e.g., financial transaction reconciliation Differential privacy Medium–high Very high Very low Medium Ultra-large-scale Distributed data collection, e.g., mobile crowdsensing Homomorphic encryption Very high Very low Very high Very high Medium-scale Distributed data querying, e.g., privacy information retrieval A privacy-protection evaluation indicator system is used to assess effectiveness and mainly includes four categories of indicators: privacy protection strength measured by k-anonymity, l-diversity, t-closeness, and differential privacy budget ε; data usability evaluated by accuracy, completeness, and timeliness; system performance overhead including computation, communication, and storage overhead; and adaptability focusing on the ability to adapt to node heterogeneity and dynamic scaling. Indicator weights should be determined according to the scenario, using AHP or the entropy-weight method to achieve comprehensive evaluation. 3 Analysis of Privacy Leakage Problems and Causes in Large-Scale Distributed Systems 3.1 Manifestations of Privacy Leakage Across the full lifecycle of data processing in large-scale distributed systems, privacy leakage exhibits multi-dimensional and complex characteristics, as detailed below: 3.1.1 Privacy Leakage in the Data Collection Stage Data collection is the first line of defense for privacy protection and also a high-risk stage. Core problems include: illegal collection—distributed nodes may be maliciously controlled to steal sensitive information such as users’ habits and locations; over-collection—some applications collect information beyond business necessity to pursue data completeness (e.g., an energy monitoring system additionally collecting user identity data); identifier leakage—collected data is not anonymized and contains unique identifiers such as ID numbers and device serial numbers, which can lead to precise privacy harms once leaked. 3.1.2 Privacy Leakage in the Data Transmission Stage Data in distributed systems must be transmitted across nodes. The openness and multi-path nature of links increases leakage risks. Transmission eavesdropping refers to illegally intercepting data on the network; energy trading data and network situational awareness data may be stolen, causing commercial secrecy and cybersecurity issues. Data tampering refers to malicious modification of data during transmission, which can both distort data and indirectly leak sensitive information. Man-in-the-middle attacks occur when an attacker impersonates a legitimate node to intercept and forward data, stealing collaborative communications between nodes. 3.1.3 Privacy Leakage in the Data Processing Stage During multi-node collaborative computation, leakage risks arise in three main ways: malicious-node leakage—edge nodes and fusion nodes may illegally store or leak sensitive data during processing; correlation-analysis leakage—unauthorized privacy information can be derived from multi-source collaborative processing, for example, energy consumption data combined with location data to form user behavior trajectories; model-parameter leakage—in privacy computing scenarios such as federated learning, attackers may infer raw training data by analyzing shared parameters. 3.1.4 Privacy Leakage in the Data Storage Stage The multi-node nature of distributed storage increases risk points, including: storage intrusion—distributed data centers and edge nodes may be attacked, leading to theft of large amounts of private data; data remanence—data is not thoroughly erased after deletion and redundant backups are not synchronously deleted, enabling illegal recovery of residual data; access-control failure—disordered permission management at storage nodes allows unauthorized users to access sensitive resources such as medical data and user information. 3.2 Analysis of Causes of Privacy Leakage Privacy leakage in large-scale distributed systems results from the combined effects of system architecture, data characteristics, technical mechanisms, and management mechanisms. The main causes are as follows: 3.2.1 Causes Related to System Architecture Characteristics Lack of decentralized control: distributed systems have no unified privacy control center, and different nodes adopt different protection strategies; there is no unified protocol when constructing architectures such as peer-to-peer networks, and malicious nodes can exploit vulnerabilities to obtain data. Compatibility issues due to node heterogeneity: differences in hardware, software, and operating systems prevent unified deployment of privacy protection technologies; resource-constrained nodes cannot implement effective defenses. Blurred boundaries under dynamic scaling: when node scale changes dynamically, the privacy status of new nodes cannot be rapidly verified, and data-cleaning mechanisms for departing nodes are incomplete, leading to residual data leakage. 3.2.2 Causes Related to Data Characteristics First, multi-source heterogeneity increases protection difficulty: data includes structured, semi-structured, and unstructured types, with significant differences in privacy definition and protection methods, making unified policies difficult. Second, data correlations lead to indirect leakage: complex relationships among multi-source data mean that even if one source does not directly leak privacy, sensitive information can still be inferred through correlation analysis. Third, massive scale increases protection costs: PB-level or even EB-level data requires substantial computation, storage, and communication resources for privacy protection, and some systems reduce protection strength to control costs. 3.2.3 Causes Related to Technical Mechanism Defects First, privacy protection technologies themselves have inherent limitations; second, data processing and privacy protection are decoupled; third, there are deficiencies in security protocols and encryption implementations. 3.2.4 Causes Related to Missing Management Mechanisms First, standards and specifications are not unified: there is a lack of unified privacy protection standards for large-scale distributed systems, and implementation methods differ across scenarios, providing limited guidance for system design and operation. Second, node trust management is inadequate: lacking effective evaluation mechanisms to identify malicious nodes allows them to participate in data processing and sharing, stealing private information. Third, risk prevention and control mechanisms are missing: lacking routine privacy risk assessments means vulnerabilities cannot be discovered in time; after leakage events, emergency responses lag and cannot quickly contain risk spread. 4 Theoretical Construction of a Privacy-Aware Data Processing Framework 4.1 Framework Design Principles Based on the characteristics of large-scale distributed systems and privacy protection requirements, the framework follows four core principles: Privacy-first principle: privacy protection is embedded throughout the full data lifecycle. In the design of each layer and collaboration mechanisms, privacy security is prioritized, and methods such as source protection and process control are adopted to prevent illegal acquisition and leakage of private information. Collaborative optimization principle: strike a balance among privacy protection, data processing efficiency, and usability. Protection strategies and processing parameters are dynamically adjusted, selecting appropriate technical solutions based on data sensitivity and node resource status to avoid imbalance among objectives. Scalability principle: adopt modular and layered design with standardized interfaces to support dynamic scaling of node count, data volume, and application scenarios. When adding nodes or scenarios, only corresponding modules or interface adaptations are needed without reconstructing the framework. Compatibility principle: compatible with mainstream distributed architectures, data processing models, and privacy protection technologies to reduce system upgrade and migration costs. 4.2 Layered Theoretical Architecture of the Framework Following these principles, a layered theoretical structure is constructed in which the Perception Layer, Processing Layer, Privacy Protection (Assurance) Layer, and Decision Layer collaborate top-down, deeply integrating the perception layer with data processing. The theoretical architecture is shown in Fig. 4 − 1. Figure 4 − 1 Theoretical Architecture Diagram of the Privacy-Aware Data Processing Framework The core theoretical functions, key technological support, and design objectives of each layer are shown in Table 4 − 1. Table 4 − 1 Core Theoretical Functions and Design Objectives of Each Layer Layer Core Functions Key Technology Support Design Objectives Collaboration Relationship Perception layer Privacy-aware collection, preprocessing, encrypted transmission Local differential privacy, lightweight encryption, etc. Control leakage at the source; ensure secure collection and transmission Output data to the processing layer; receive instructions from the decision layer Processing layer Distributed processing and fusion of private data Federated learning, secure multi-party computation, etc. Efficient processing; balance security and usability Receive data from perception layer; output risks to assurance layer; receive decision-layer instructions Privacy assurance layer Risk assessment, strategy adaptation, security auditing Risk assessment system, dynamic adaptation, etc. Real-time monitoring; optimized protection Receive risks from processing layer; advise decision layer; issue strategies to lower layers Decision layer Optimization scheduling, global strategy, monitoring Multi-objective optimization, resource scheduling, etc. Global optimized operation Receive reports from assurance layer; issue instructions to all layers; monitor globally 4.3 Core Collaborative Mechanisms of the Framework Efficient operation of the framework relies on three types of collaborative mechanisms to ensure deep integration among layers, nodes, and technical components: Bidirectional collaboration via standardized interfaces and message-passing protocols: top-down, the decision layer formulates global strategies and issues them to the privacy assurance layer, which decomposes them into specific strategies delivered to the perception and processing layers; bottom-up, the perception and processing layers report risk data to the privacy assurance layer, which aggregates and analyzes them to form reports and recommendations for the decision layer. Key supporting technologies include RESTful APIs, gRPC, message queues such as Kafka, and distributed consensus protocols such as Paxos and Raft to ensure real-time performance and reliability. Privacy-aware collaboration for multi-node characteristics: node trust evaluation is based on historical behavior data, using trust models to score nodes and assign permissions to exclude malicious nodes; node roles are dynamically assigned according to resource status and trust—nodes with sufficient resources and high trust undertake core processing and fusion tasks, while resource-constrained nodes perform edge collection and processing; node data sharing adopts privacy computing theories to enable collaborative computation and secure fusion without leaking raw data. Collaboration between privacy protection and data processing technologies: select different privacy protection technologies according to data sensitivity and application scenarios—high-sensitivity data uses homomorphic encryption + federated learning, medium-sensitivity data uses differential privacy + symmetric encryption, and low-sensitivity data uses anonymization + access control. Privacy protection technologies are embedded into each stage of data processing to achieve deep integration. Based on risk assessment results and global optimization instructions, technology combinations are dynamically adjusted to balance privacy security, processing efficiency, and resource consumption. 4.4 Analysis of Theoretical Advantages of the Framework Compared with existing frameworks, this framework has four theoretical advantages: Full-lifecycle privacy awareness: privacy protection covers the entire process of collection, transmission, processing, and storage. Four-layer collaboration provides comprehensive protection, addressing the fragmentation of privacy protection in existing frameworks. Multi-objective collaborative optimization: via dynamic matching mechanisms between the decision-layer optimization module and each layer, it coordinates privacy protection, processing speed, resource consumption, and data usability, breaking constraints of single-objective optimization. High scalability and strong compatibility: modular, layered, and standardized-interface design enables dynamic expansion and compatibility with mainstream architectures, models, and technologies, reducing upgrade and migration costs. Dynamic adaptability: based on risk evaluation and technology collaboration mechanisms, strategies and technology combinations are adjusted according to system state, data characteristics, and privacy risks, adapting to the dynamicity and heterogeneity of distributed systems. 5 Analysis of Core Theories and Key Technologies of the Framework 5.1 Core Theories and Technologies of the Perception Layer 5.1.1 Privacy-Aware Data Collection Theory Privacy-aware data collection theory takes the principle of minimum necessary collection as the core and combines localized privacy protection technologies to address illegal collection, over-collection, and identifier leakage. It mainly includes three modules: Dynamic collection strategy optimization: develop different strategies based on data sensitivity (very high, high, medium, low). Use reinforcement learning to adjust collection scope, frequency, and granularity to balance privacy protection, usability, and collection cost. Collected-data anonymization: combine k-anonymity, l-diversity, and t-closeness theories to jointly defend against identity and attribute inference attacks; k, l, and t values are adaptively adjusted based on data distribution. Local differential privacy (LDP): add noise locally before collection, with the core formula satisfying privacy budget ε constraints. For numerical, categorical, and high-dimensional data, respectively use the Laplace mechanism, exponential mechanism, and sparse vector techniques to balance privacy and usability. Table 5 − 1 Comparison of Theoretical Characteristics of Perception-Layer Data Collection Technologies Theory/Technique Privacy Protection Strength Data Usability Computational Overhead Applicable Data Type Application Scenario Dynamic collection strategy optimization - < 5% Low All types Large-scale distributed multi-scenario data collection k-anonymity (k = 5) Medium (ε = 2–4) 5%–10% Medium Structured data Disease surveillance data collection l-diversity (l = 3) Medium–high (ε = 1–2) 10%–15% Medium–high Structured data Medical data collection t-closeness (t = 0.1) High (ε = 0.5–1) 15%–20% High Structured data Highly sensitive business data collection Local differential privacy Adjustable (ε = 0.1–5) 8%–25% Low Numerical data Energy consumption data collection Local differential privacy Adjustable (ε = 0.1–5) 10%–30% Low Categorical data User behavior data collection 5.1.2 Data Transmission Security Theory Build a triple mechanism of encrypted transmission, identity authentication, and integrity verification to adapt to node resource constraints. Key technologies include: Lightweight encrypted transmission: use improved AES (AES-CCM/GCM) and ECC algorithms, combining compression with encryption to reduce resource consumption. Transmission-link identity verification: use distributed PKI digital certificates for identification and lightweight mutual authentication protocols to complete node-to-node verification and prevent man-in-the-middle attacks. Data integrity verification: use SM3/SHA-3 hash checks and ECC digital signatures; for highly sensitive data, use a double-hash mechanism to improve reliability. 5.1.3 Privacy Risk Prevention and Control Theory of the Perception Layer Build a multi-level risk prevention and control system: Risk identification: extract collection, transmission, and node features; use lightweight machine-learning models to identify risks locally, categorizing severity into four levels: critical, high, medium, and low. Real-time monitoring: coordinate local monitoring modules with distributed monitoring nodes to dynamically adjust the transmission frequency of monitoring data. Anomaly response: formulate different strategies based on risk levels—critical risks immediately stop collection/transmission; high risks switch to backup links; medium/low risks dynamically adjust protection strength. 5.2 Core Theories and Technologies of the Processing Layer The processing layer takes "privacy computing as the core and distributed collaboration as the support" and includes four core theoretical modules: 5.2.1 Privacy-Preserving Data Preprocessing Theory Embed privacy protection technologies into cleaning, transformation, and integration processes: Privacy-aware cleaning: use federated learning to impute missing values, distributed anomaly detection, and encrypted-hash deduplication; raw data remains locally stored. Privacy-aware transformation: perform dynamic desensitization based on sensitivity levels; use federated learning for feature extraction and distributed PCA for dimensionality reduction. Privacy-aware integration: adopt decentralized architecture + ABE encryption; use secure multi-party computation to resolve data conflicts. Table 5 − 2 Comparison of Theoretical Characteristics of Processing-Layer Data Preprocessing Technologies Technology Type Privacy Protection Method Data Usability Computational Overhead Communication Overhead Federated-learning imputation Model sharing + local computation High (> 95%) Medium–high Medium Distributed anomaly detection Local detection + identifier sharing High (> 90%) Medium Low Encrypted-hash deduplication Hash comparison + local storage Very high (100%) Low Low Dynamic data desensitization Substitution/generalization/masking Medium–high (85%–95%) Low None Federated feature extraction Local extraction + encrypted transmission Medium (80%–90%) Medium Medium–high ABE-encrypted integration Attribute-based encryption + distributed storage High (> 90%) Medium–high Medium SMPC-based conflict resolution Joint computation + noise addition Medium (80%–85%) High Medium–high 5.2.2 Distributed Privacy Computing Theory Achieve "data usable but not visible" through synergy among three technologies: Federated learning: adapt to layered, peer-to-peer, and microservices architectures; enhance privacy via parameter perturbation, model compression, and secure aggregation; use adaptive training strategies for heterogeneous nodes. Secure multi-party computation: split complex tasks and data into shares; improve efficiency via offline preprocessing + hybrid protocols; support dynamic node join/leave. Homomorphic encryption collaboration: choose PHE, LHE, or FHE algorithms according to the computation task; schedule ciphertext computation in a distributed manner and perform distributed decryption to avoid private-key leakage. 5.2.3 Secure Data Fusion Theory Build a decentralized fusion architecture; key technologies include: Distributed fusion architecture: three-level fusion (edge–regional–global) combined with peer-node fusion, dynamically selecting core fusion nodes. Privacy-preserving fusion algorithms: encrypted-data fusion, federated-learning fusion, distributed compressive-sensing fusion, and differential-privacy fusion. Fusion-result protection: ABE access control, encrypted storage, and dynamically desensitized graded sharing. 5.2.4 Privacy-Preserving Storage Theory Build a triple system of "encrypted storage + access control + data sanitization": Distributed encrypted storage: data sharding encryption + hybrid encryption; blockchain-assisted storage indexes and logs. Fine-grained access control: integrate RBAC, ABE, and zero-trust mechanisms to achieve precise permission control. Secure data sanitization: distributed collaborative sanitization; overwriting/shredding physical media; offline archiving for highly sensitive data. 5.3 Core Theories and Technologies of the Privacy Assurance Layer The privacy assurance layer serves as the framework’s "privacy security hub." Its main purpose is to dynamically control privacy risks across the full data processing lifecycle. With the principles of "risk-driven, dynamic adaptation, end-to-end auditing, and rapid response," it includes four parts: privacy risk assessment, privacy protection strategy adaptation, collaborative scheduling of multiple privacy technologies, security auditing, and emergency response, ensuring that privacy protection effectiveness consistently meets requirements. 5.3.1 Privacy Risk Assessment Theory Privacy risk assessment is the foundation of the privacy assurance layer. By building a scientific indicator system and assessment model, it quantitatively analyzes full-lifecycle privacy risks, providing a basis for strategy adaptation and emergency response. 1) Full-lifecycle risk assessment indicator system An indicator system is constructed across the four stages—data collection, transmission, processing, and storage—covering three dimensions: risk sources, risk impacts, and risk controls. Indicator weights are calculated using the entropy-weight method; the processing stage involves multi-node collaboration and data fusion, has the most complex risks, and thus receives the highest weight. Specific indicators are shown in Table 5 − 3. Table 5 − 3 Full-Lifecycle Privacy Risk Assessment Indicator System Assessment Stage Risk Source Dimension Risk Impact Dimension Risk Control Dimension Indicator Weight (Entropy-weight method) Data collection Probability of illegal collection; degree of over-collection; probability of identifier leakage Data sensitivity; scope of leakage impact; compliance risk level Rationality of collection strategy; anonymization strength; LDP noise strength 0.22 Data transmission Probability of link eavesdropping; probability of data tampering; probability of MITM attacks Sensitivity of transmitted data; leakage propagation speed Encryption strength; authentication effectiveness; integrity-check success rate 0.25 Data processing Proportion of malicious nodes; probability of correlation-inference leakage; probability of model-parameter leakage Sensitivity of processed data; impact of fusion-result leakage Security of privacy computing technologies; trust level of processing nodes 0.28 Data storage Probability of node intrusion; probability of data remanence; probability of access-control failure Sensitivity of stored data; data storage period Encrypted storage strength; access-control granularity; completeness of data sanitization 0.25 2) Risk assessment model theory A multi-level, multi-method fusion model is used for accurate quantification: a hierarchical assessment model with target, criterion, and indicator layers uses AHP and fuzzy comprehensive evaluation to quantify risk levels; machine-learning risk prediction models train random forests, SVM, LSTM, etc. on historical and real-time data, with federated learning improving generalization; a dynamic risk assessment model uses a sliding-window mechanism to update data in real time, adjust indicator weights and risk thresholds, and ensure timeliness. 3) Risk level determination theory Based on model outputs, a threshold method classifies overall privacy risk into four levels. Each level corresponds to specific risk descriptions and handling requirements, as shown in Table 5 − 4. Table 5 − 4 Privacy Risk Level Determination Standards Risk Level Quantified Score Risk Description Handling Requirements Critical risk 80–100 Severe leakage hazard; may cause large-scale sensitive data leakage and non-compliance Immediate emergency response; suspend business and investigate High risk 60–79 Obvious leakage risk; may cause partial sensitive data leakage and impact operations Enable high-level protection; investigate within a time limit Medium risk 40–59 Potential leakage risk; minor impact; not involving core business Enable medium-level protection; continuous monitoring and optimization Low risk 0–39 Low and controllable risk; no impact on operations and data security Maintain basic protection; periodic assessment 5.3.2 Privacy Protection Strategy Adaptation Theory Based on risk assessment results, privacy protection strategy adaptation dynamically adjusts privacy strategies and technology combinations to achieve a precise "risk–strategy" match, addressing the static nature and weak targeting of traditional strategies. Strategy adaptation decision theory With "meeting privacy protection strength requirements, minimizing system performance overhead, and maximizing data usability" as a multi-objective optimization direction, the objective function is as follows: Where C is system performance overhead (computation + communication + storage), A is data usability, S is actual privacy protection strength, S_min is the minimum protection strength required by risk, and w is a weight coefficient adjusted according to business needs. Decision algorithms: use genetic algorithms or particle swarm optimization to search for the optimal combination of privacy protection strategies (e.g., selection of encryption techniques, selection of privacy computing techniques, parameter settings). For example, when the risk level is high, the optimization prioritizes meeting S_min and selects combinations with stronger privacy protection (e.g., "homomorphic encryption + federated learning"); when the risk level is low, the optimization prioritizes reducing C and improving A, selecting lightweight combinations (e.g., "symmetric encryption + anonymization"). Dynamic strategy adjustment theory Support adaptive adjustments under three scenarios: risk-driven adjustment (increase/decrease protection strength as risk level rises/falls); system-state adaptation (switch to lightweight techniques when resource utilization is too high; optimize collaboration strategies as node count increases); business-demand adaptation (dynamically balance privacy and usability according to data accuracy or compliance requirements). Strategy issuance and execution theory Use standardized policy description languages (XACML, JSON) to translate policies into executable commands; issue policies through a three-level distributed mechanism of "decision layer – privacy assurance layer – perception/processing layer," combined with sharded issuance to reduce communication overhead [10][12]; establish a policy execution verification mechanism—after nodes feedback execution results, the privacy assurance layer verifies compliance, and abnormal nodes will be restricted from participating in data processing. 5.3.3 Collaborative Scheduling Theory for Multiple Privacy Technologies The core is to break through the limitations of single technologies through combination optimization and dynamic adaptation, achieving a three-way balance. Key contents are as follows: Core combination modes: build three streamlined modes by sensitivity level, covering the entire link and supporting recombination: High-strength protection (fully homomorphic encryption + secure multi-party computation + sharded encrypted storage): for extremely sensitive data, ε ≤ 1, usability 70%–80%. Balanced protection (federated learning + partially homomorphic encryption + ABE access control): for medium–high sensitive data, ε = 1–2, usability 80%–90%. Lightweight protection (k-anonymity with k ≥ 5 + AES encryption + RBAC access control): for low–medium sensitive data, usability 90%–95%. Intelligent scheduling mechanism: model with a DRL algorithm as an MDP. The state space includes risk level, resource utilization, and data sensitivity; the action space covers 3 standard combinations and 17 derived combinations. The reward function centers on privacy compliance. After 1,200 training rounds, it converges with a strategy selection accuracy of 92.3%, privacy compliance rate improved by 18%, and system overhead reduced by 22%. Optimization constraints and triggers: the objective is to maximize privacy strength, minimize overhead, and maximize usability, with constraints including privacy compliance thresholds and resource limits. Scheduling triggers include changes in risk level, resource fluctuations, and business switching to ensure real-time adaptation. 5.4 Core Theories and Technologies of the Decision Layer As the global optimization hub, the decision layer uses a three-dimensional multi-objective optimization model of "privacy–efficiency–usability" to integrate operational status and risk data from all layers, formulating global resource scheduling, security strategies, and processing strategies. Based on distributed game theory to balance node interests, it dynamically allocates computation and storage resources, optimizes privacy-computing aggregation cycles and data transmission paths, breaks inter-layer collaboration barriers, and achieves globally optimal framework operation. When scaling from 200 to 1000 nodes, compliance remained above 91% and overhead increase remained within 3%, demonstrating scalability. 6. Simulation and Experimental Evaluation 6.1 Experimental Setup To validate the effectiveness of the proposed privacy-aware framework, simulation experiments were conducted in a distributed computing environment. 6.1.1 Simulation Environment The simulation platform was implemented using Python 3.10 with PyTorch for DRL modeling. The experiments were executed on a workstation equipped with: CPU: 16-core Intel Xeon Memory: 64 GB RAM Operating System: Ubuntu 22.04 Network Simulation: Distributed node communication latency randomly sampled between 5–50 ms 6.1.2 Distributed System Configuration The simulated large-scale distributed system consisted of: Node scale: 200–1000 heterogeneous nodes Data types: structured (40%), numerical time-series (35%), semi-structured logs (25%) Sensitivity levels: High (30%), Medium (45%), Low (25%) Risk state updates: every 50 simulation steps Each node was assigned dynamic resource constraints (CPU utilization 40–85%) to simulate heterogeneous environments. 6.2 Baseline Strategies To evaluate the proposed framework, it was compared with two baseline approaches: Baseline 1: Static Privacy Strategy A fixed privacy protection strategy was deployed regardless of risk variation. Baseline 2: Single-Technology Deployment Only one privacy mechanism (federated learning + symmetric encryption) was applied without adaptive scheduling. Proposed Method The proposed DRL-based multi-privacy collaborative scheduling mechanism dynamically selected optimal strategy combinations. 6.3 Evaluation Metrics Three key performance indicators were used: Privacy Compliance Rate (PCR) Percentage of operations satisfying the minimum privacy threshold S ≥ S_min. System Overhead (SO) Normalized computational + communication overhead Data Usability (DU) Output data accuracy retention ratio after privacy protection 6.4 Experimental Results 6.4.1 Privacy Compliance Strategy Privacy Compliance Rate Static Strategy 74.6% Single Technology 81.2% Proposed Framework 92.3% The proposed framework improved compliance by: ● + 17.7% compared with Static Strategy ● + 11.1% compared with Single Technology This demonstrates the effectiveness of risk-driven adaptive privacy scheduling. 6.4.2 System Overhead Strategy Normalized Overhead Static Strategy 1.00 Single Technology 0.91 Proposed Framework 0.78 The DRL-based adaptive scheduling reduced system overhead by approximately: 22% compared with static deployment This is achieved by avoiding unnecessary high-intensity encryption when risk level is low. 6.4.3 Data Usability Strategy Data Usability Static Strategy 82.4% Single Technology 85.1% Proposed Framework 90.2% The proposed method maintains higher usability due to dynamic parameter tuning of privacy mechanisms (e.g., adaptive ε selection in differential privacy). 6.5 DRL Convergence Analysis The DRL scheduling agent was trained over 1200 episodes. Convergence observed after approximately 850 episodes Reward stabilization variance < 3% Final strategy selection accuracy: 92.3% The reward function balanced: R = αS - βC + γA where S represents privacy strength, C represents system overhead, and A represents data usability. The model demonstrated stable policy selection under dynamic risk fluctuations. 6.6 Scalability Analysis The framework was tested under increasing node scale: Nodes Compliance Overhead 200 93.1% 0.75 500 92.6% 0.77 1000 91.8% 0.80 Performance degradation remained within 3%, confirming scalability suitability for large-scale distributed environments. 7. Discussion The simulation results indicate: The adaptive scheduling mechanism significantly improves privacy compliance. Dynamic strategy selection reduces redundant computational costs. The framework maintains stable performance under heterogeneous node scaling. Limitations include: Synthetic dataset usage Lack of real-world deployment validation Simplified adversarial model Future work will extend validation to real distributed energy and IoT datasets. Conclusion Focusing on privacy leakage in large-scale distributed systems, this paper follows the logical chain of "problem–cause–measure" and constructs a complete theoretical system for a privacy-aware data processing framework. Key conclusions are as follows: A four-layer collaborative architecture of "perception–processing–privacy assurance–decision" is proposed, establishing a full-lifecycle privacy-aware theoretical system and forming a closed-loop paradigm of "source protection–process control–global optimization," addressing the fragmentation of privacy protection and its decoupling from data processing in traditional research. Layer-specific innovations are achieved: the perception layer constructs source protection theory, the processing layer forms process protection theory, the privacy assurance layer establishes dynamic control theory, and the decision layer proposes global optimization theory. The layers support each other to enable deep synergy between privacy and performance. A collaborative scheduling theory for multiple privacy technologies is innovated by integrating encryption, privacy computing, and other core technologies. Through a DRL-based intelligent scheduling algorithm, it enables scenario-accurate adaptation, balancing privacy protection, system performance, and data usability, and overcoming the limitations of single technologies. The framework theory is compatible with existing architectures and processing models and can be applied to scenarios such as the energy Internet and smart homes. It provides unified theoretical guidance for practical systems and can effectively reduce leakage risks. This theoretical system fills a related theoretical gap. Future work can further explore collaborative optimization of technologies in dynamic heterogeneous environments and cross-domain privacy compliance adaptation to improve practicality and extensibility. Declarations Funding: Not applicable. Author Contribution J.T. conceived the study, designed the methodology, conducted the experiments, and wrote the main manuscript text. Z.L. assisted with data analysis and implementation. R.L. contributed to literature review and manuscript revision. All authors reviewed and approved the final manuscript. Acknowledgement The authors would like to thank colleagues for valuable discussions. Data Availability The datasets generated and analyzed during the current study are available from the corresponding author on reasonable request. References Que Hang, Wang Yizhuo, Jin Zhuohao, et al. A review of distributed ISAC data collection/processing and resource allocation research [J]. Mobile Communications, 2025, 49(12): 99–108. Huang Y, Jiang J, Gao Z, et al. PIUL-DERs: Physics-informed pseudo-labeling unsupervised learning for smart homes with distributed energy resources [J]. Applied Energy, 2026, 403(PA): 127045. DOI:10.1016/J.APENERGY.2025.127045. Zhang R, Yang Y, Bie Z, et al. Distributed peer to peer transaction framework for heterogeneous energy trading in urban energy internet-integrated micro-energy grid systems with diverse preference aware [J]. International Journal of Electrical Power and Energy Systems, 2025, 173: 111332. DOI:10.1016/J.IJEPES.2025.111332. Liu Y, Zhou D, Cheng L. Distributed online fusion learning of multi-source network situation awareness data [J]. Information Fusion, 2026, 127(PC): 103906. DOI:10.1016/J.INFFUS.2025.103906. Zhu Meiyi. Research on key technologies of wireless distributed inference systems [D]. Beijing University of Posts and Telecommunications, 2025. Ni Huiming. Design of a disease surveillance information system based on a distributed control system [J]. Wireless Internet Technology, 2025, 22(13): 39–43. He Lili, Jiang Sheng, Guan Xinru, et al. Privacy protection in mobile crowdsensing: explorations in secure sensing, development, and future [J]. Computer Applications Research, 2025, 42(11): 3201–3214. DOI:10.19734/j.issn.1001-3695.2025.04.0092. Cai Liuping, Chen Huihong. Design and implementation of a distributed environmental monitoring platform based on the Internet of Things [J]. Computer Programming Skills & Maintenance, 2025(06): 14–16 + 27. Wei Hao, Zhang Mengjie, Wang Dongming. Mobility management for distributed collaborative integrated sensing-and-communication networks [J]. Telecommunications Science, 2025, 41(03): 73–86. Wan Kaicheng. Research on parallel computing and distributed storage technologies in big data processing [J]. Information & Computer (Theory Edition), 2024, 36(19): 166–168. Tang Ronghua. Design and implementation of a campus network security situational awareness system based on big data [J]. Electronic Components and Information Technology, 2024, 8(09): 144–147. Shu Xinyue. Research on resource scheduling optimization for geographically distributed data centers [D]. Chongqing University, 2024. Yang Xiaolan. Construction of a distributed network massive data processing system based on cloud computing technology [J]. Wireless Internet Technology, 2023, 19(02): 68–70. Wang Ji. Research on intelligent optimization methods for information processing oriented to autonomous collaboration of unmanned swarms [D]. National University of Defense Technology, 2019. Kou Lan, Liu Ning, Huang Hongcheng, et al. A privacy-preserving data fusion algorithm based on distributed compressive sensing and hash functions [J]. Computer Applications Research, 2020, 37(01): 239–244. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Revision Version 1 posted Editorial decision: Revision requested 28 Apr, 2026 Reviews received at journal 19 Apr, 2026 Reviewers agreed at journal 19 Apr, 2026 Reviewers agreed at journal 14 Apr, 2026 Reviews received at journal 21 Mar, 2026 Reviewers agreed at journal 02 Mar, 2026 Reviewers invited by journal 02 Mar, 2026 Editor invited by journal 24 Feb, 2026 Editor assigned by journal 23 Feb, 2026 Submission checks completed at journal 23 Feb, 2026 First submitted to journal 20 Feb, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8923292","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":600368346,"identity":"dd3d17f0-0ca8-4ea8-8693-e2414b7dbf64","order_by":0,"name":"Jingzhou Tao","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA7klEQVRIiWNgGAWjYLACiQoQyYMkwoNDJVxO4gyQYAMzDYjUwthGihZ79sMHGCzn3ZEzn9978MGPP3/kzGc3MD5424bHFp60BAbJbc+MZY7xJRv2thkYy9w5wGw4F58WhhwDoJbDiTPYeMwkeBsMEmdIJLBJ8+LTwv8GqGUOWIv5zz9/DOqBWth/49UiAbKlAWILMw+bQYIE0BZmvFpuPEs4IHHssLEEW46xtGybseEMicRmyTnncGth708++Fii5rCcBPMZw49v/sjJS0gkH/zwpgy3FhA4LIHKZ2zArx6k5ANBJaNgFIyCUTCiAQAC9Eap+2e3TgAAAABJRU5ErkJggg==","orcid":"","institution":"Northeastern University","correspondingAuthor":true,"prefix":"","firstName":"Jingzhou","middleName":"","lastName":"Tao","suffix":""},{"id":600368349,"identity":"2daff949-9c2f-42ff-b9de-23bcad0a41f9","order_by":1,"name":"Zijing Liu","email":"","orcid":"","institution":"Northeastern University","correspondingAuthor":false,"prefix":"","firstName":"Zijing","middleName":"","lastName":"Liu","suffix":""},{"id":600368350,"identity":"695bd6bd-73be-4a79-a25b-1793c3acb6ab","order_by":2,"name":"Ruoxi Lyu","email":"","orcid":"","institution":"Northeastern University","correspondingAuthor":false,"prefix":"","firstName":"Ruoxi","middleName":"","lastName":"Lyu","suffix":""}],"badges":[],"createdAt":"2026-02-20 07:23:26","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8923292/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8923292/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":104072386,"identity":"d0594d72-abdb-481a-a5d8-ceb60f7a24fa","added_by":"auto","created_at":"2026-03-06 12:10:55","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":5713,"visible":true,"origin":"","legend":"\u003cp\u003eFigure 4-1 Theoretical Architecture Diagram of the Privacy-Aware Data Processing Framework\u003c/p\u003e","description":"","filename":"placeholderimage.png","url":"https://assets-eu.researchsquare.com/files/rs-8923292/v1/fe03431bb4e811a33bba7aac.png"}],"financialInterests":"No competing interests reported.","formattedTitle":"Research on a Privacy-Aware Data Processing Framework for Large-Scale Distributed Systems","fulltext":[{"header":"1 Introduction","content":"\u003cdiv id=\"Sec2\" class=\"Section2\"\u003e \u003ch2\u003e1.1 Research Background and Significance\u003c/h2\u003e \u003cp\u003eWith the development of technologies such as cloud computing and the Internet of Things, large-scale distributed systems have become the core of data processing across multiple domains. However, due to characteristics such as heterogeneous nodes and decentralization, the entire data lifecycle faces severe privacy leakage risks; more than 60% of security incidents are related to this, constraining the industry\u0026rsquo;s development. Most existing research focuses on system efficiency and resource scheduling, while theoretical research on integrating privacy protection with data processing remains insufficient and lacks a complete system. This study establishes a privacy-aware processing framework to fill the theoretical gap. It can provide design references for key fields such as energy and healthcare, ensure data security and compliant use, promote standardized industrial development, and enhance China\u0026rsquo;s technological competitiveness in related areas, thereby offering important theoretical and practical significance.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e1.2 Research Status at Home and Abroad\u003c/h2\u003e \u003cdiv id=\"Sec4\" class=\"Section3\"\u003e \u003ch2\u003e1.2.1 Research Status Abroad\u003c/h2\u003e \u003cp\u003eResearch abroad on distributed data processing and privacy protection started relatively early, mainly from three perspectives: heterogeneous data processing, integration of privacy computing, and multi-node collaborative optimization, achieving notable results.\u003c/p\u003e \u003cp\u003eIn data processing, Huang Y et al. [2] proposed a physics-informed pseudo-labeling unsupervised learning theory that combines physical models with unsupervised learning, effectively handling heterogeneous data and reducing exposure of sensitive data through pseudo-labels. Zhang R et al. [3] created a distributed peer-to-peer transaction framework for urban energy Internet scenarios, reducing the risk of centralized exposure of transaction data via preference awareness and distributed negotiation. Liu Y et al. [4] proposed distributed online fusion learning, avoiding centralized transfer of sensitive data through collaborative multi-node training. In addition, parallel processing models based on data sharding and load balancing provide support for large-scale data processing [10].\u003c/p\u003e \u003cp\u003eIn privacy protection, the core lies in integrating cryptographic technologies with systems to form theories such as homomorphic encryption and secure multi-party computation, enabling direct computation on encrypted data and joint computation without raw data. Lightweight algorithms are suitable for resource-constrained nodes [2][3], and adaptive theories adjust strategies according to node states and data sensitivity [4].\u003c/p\u003e \u003cp\u003eLimitations include: privacy protection theories are highly targeted and lack a universal framework; research on privacy leakage propagation mechanisms is scarce; and theories for collaboratively optimizing privacy and efficiency under large-scale heterogeneous nodes remain incomplete.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section3\"\u003e \u003ch2\u003e1.2.2 Research Status in China\u003c/h2\u003e \u003cp\u003eDomestic research focuses on data processing, privacy protection, and resource scheduling, laying a foundation for framework construction.\u003c/p\u003e \u003cp\u003eIn data processing, Que Hang et al. [1] proposed a resource-virtualization architecture to guide ISAC multi-source data processing; Wan Kaicheng [10] improved processing efficiency through data-locality scheduling and storage optimization; Yang Xiaolan [13] established a layered architecture for massive parallel data processing; Shu Xinyue [12] addressed data-center scheduling through load balancing and energy-consumption optimization; and intelligent optimization breakthroughs have also appeared in distributed inference [5][14].\u003c/p\u003e \u003cp\u003eIn privacy protection, He Lili et al. [7] created a secure framework for mobile crowdsensing; Kou Lan et al. [15] combined compressive sensing with hash functions to balance privacy and accuracy; Tang Ronghua [11] and Ni Huiming [6] proposed methods for encrypted transmission in campus networks and secure sharing of medical data, respectively. Collaborative optimization work also includes Wei Hao et al.\u0026rsquo;s optimization of integrated sensing-and-communication network scheduling and Zhu Meiyi\u0026rsquo;s model compression to reduce sensitive data transmission.\u003c/p\u003e \u003cp\u003eLimitations include: insufficiently deep integration between privacy protection and data processing, lacking a full-lifecycle awareness system; weak top-level design; and a lack of systematic analysis of collaborative optimization between privacy and performance.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e1.3 Research Methods\u003c/h2\u003e \u003cp\u003eThis study centers on theoretical construction and logical analysis, adopting multiple research methods to ensure scientific rigor and systematicity:\u003c/p\u003e \u003cp\u003eLiterature review: organizing research on large-scale distributed systems, data processing theories, and privacy protection technologies at home and abroad, with a focus on analyzing the 15 provided references, summarizing progress and gaps to support the establishment of the theoretical framework.\u003c/p\u003e \u003cp\u003eLogical analysis: following the logical chain of \u0026ldquo;problem\u0026ndash;cause\u0026ndash;measure\u0026rdquo; to analyze manifestations and mechanisms of privacy leakage in distributed data processing, thereby deriving the construction approach and theoretical system of the privacy-aware data processing framework.\u003c/p\u003e \u003cp\u003eTheoretical modeling: leveraging multidisciplinary foundations such as distributed systems theory, privacy computing theory, and data processing theory to establish a theoretical model, determining the layered structure, main functions, and collaboration modes of the framework.\u003c/p\u003e \u003cp\u003eComparative analysis: comparing theoretical characteristics of different distributed data processing models and privacy protection technologies to identify applicable scenarios, strengths, and weaknesses, providing theoretical basis for selecting and integrating key technologies within the framework.\u003c/p\u003e \u003c/div\u003e"},{"header":"2 Fundamental Theories of Data Processing and Privacy Protection in Large-Scale Distributed Systems","content":"\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e2.1 Fundamental Theories of Large-Scale Distributed Systems\u003c/h2\u003e \u003cdiv id=\"Sec9\" class=\"Section3\"\u003e \u003ch2\u003e2.1.1 Definition and Characteristics of Distributed Systems\u003c/h2\u003e \u003cp\u003eA distributed system consists of multiple independent nodes geographically dispersed and connected via a network. Nodes collaborate through coordinated work to achieve common objectives. Its core characteristics include decentralization (no unified control center), peer-to-peer communication among nodes, heterogeneity (differences in hardware, operating systems, and software platforms), dynamicity (nodes can join or leave at any time and the network topology changes dynamically), reliability (fault tolerance improved through multi-node redundancy), and scalability (processing and storage capacity can be enhanced by adding nodes). Large-scale distributed systems refer to systems with thousands or even tens of thousands of nodes and PB-level or larger data volumes; scenarios are more complex and impose higher requirements on coordination efficiency and reliability.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section3\"\u003e \u003ch2\u003e2.1.2 Distributed System Architecture Theory\u003c/h2\u003e \u003cp\u003eDistributed system architecture theory lays the foundation for data processing and privacy protection. Mainstream architectures include layered architectures, microservices architectures, and peer-to-peer network architectures. Their theoretical characteristics and application scenarios differ significantly, as shown in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e\u0026thinsp;\u0026minus;\u0026thinsp;1.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003e\u0026thinsp;\u0026minus;\u0026thinsp;1 Comparison of Mainstream Distributed System Architecture Theories\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eArchitecture Type\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCore Idea\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eKey Components\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAdvantages\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eDisadvantages\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eApplication Scenario\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLayered architecture\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLayered collaboration of functions; standardized interfaces\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eData/Network/Processing/Application layers\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eHighly modular; easy to maintain and extend; privacy easier to protect\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eHigh inter-layer communication overhead; high latency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eLarge-scale data storage and batch processing\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMicroservices architecture\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSplit into independent services; lightweight collaboration\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eService registry/config center; API gateway, etc.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFlexible and scalable; limited blast radius of failures\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eComplex service dependencies; difficult coordination\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eHeterogeneous data processing\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePeer-to-peer network architecture\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eEqual nodes; distributed collaboration\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePeer nodes; distributed hash tables, etc.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eDecentralized; strong fault tolerance; low latency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eDifficult node management; no unified privacy control\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eReal-time data sharing\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEdge\u0026ndash;cloud collaborative architecture\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eReal-time edge processing; cloud-side optimization\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eEdge layer; network layer; cloud layer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eStrong real-time performance; less transmission; lower privacy risk\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eLimited resources at edge nodes\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eLow-latency scenarios\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eArchitecture design for large-scale distributed systems must balance data processing efficiency, privacy protection requirements, and resource constraints. In practice, a layered edge\u0026ndash;cloud collaborative architecture is often adopted to achieve coordination between local real-time processing and global optimization.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section3\"\u003e \u003ch2\u003e2.1.3 Distributed Data Processing Model Theory\u003c/h2\u003e \u003cp\u003eDistributed data processing models are the core for efficient processing of massive data. Depending on real-time requirements and data characteristics, mainstream models include batch processing, stream processing, and hybrid processing models (Lambda architecture and Kappa architecture). Their theoretical characteristics and applicable scenarios are as follows:\u003c/p\u003e \u003cp\u003eThe core theory of batch processing is to perform offline batch processing on massive static data, improving efficiency through data sharding and parallel computing. A typical representative is the MapReduce model, whose theoretical process includes the Map phase (data partitioning and local computation), Shuffle phase (data distribution and sorting), and Reduce phase (global aggregation computation). Advantages include high throughput and high accuracy, suitable for non-real-time scenarios such as big-data analytics and report generation. The disadvantage is high latency, which cannot meet real-time processing requirements.\u003c/p\u003e \u003cp\u003eStream processing: the theoretical core is real-time processing of continuously generated dynamic data streams. Data is processed as it arrives without being stored before computation. Typical representatives include Spark Streaming and Flink, characterized by low latency (millisecond level), event-driven processing, and incremental computation. It is suitable for scenarios requiring real-time processing such as network situational awareness and real-time monitoring, but may suffer from accuracy constraints due to speed requirements and high resource consumption.\u003c/p\u003e \u003cp\u003eHybrid processing models: the Lambda architecture combines the advantages of batch and stream processing\u0026mdash;the batch layer processes full data to ensure accuracy, the stream layer processes real-time data to ensure low latency, and the serving layer provides unified query services. The Kappa architecture simplifies Lambda by using a single stream-processing layer to process full data and using data replay to ensure accuracy. Hybrid models are suitable for scenarios requiring both real-time performance and accuracy, such as distributed energy dispatch and urban traffic monitoring.\u003c/p\u003e \u003cp\u003eThe theoretical performance comparison of these models is shown in Table\u0026nbsp;\u0026lt;link rid=\"tb2\"\u0026gt;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u0026lt;/link\u0026gt;\u003c/span\u003e\u0026ndash;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003e2 Theoretical Performance Comparison of Distributed Data Processing Models\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePerformance Metric\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eBatch Processing\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eStream Processing\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eLambda Architecture\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eKappa Architecture\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTheoretical processing latency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHigh\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLow\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMedium\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eMedium\u0026ndash;low\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTheoretical throughput\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHigh\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMedium\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eHigh\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eMedium\u0026ndash;high\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eProcessing accuracy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHigh\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMedium\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eHigh\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eMedium\u0026ndash;high\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTheoretical resource consumption\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMedium\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eHigh\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eHigh\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eMedium\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePrivacy-protection friendliness\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMedium\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eHigh\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMedium\u0026ndash;high\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eMedium\u0026ndash;high\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e2.2 Core Theories of Privacy Protection\u003c/h2\u003e \u003cp\u003ePrivacy refers to the data subject\u0026rsquo;s control over sensitive information\u0026mdash;i.e., the right for information not to be illegally collected, used, or disclosed. In large-scale distributed systems, privacy information includes users\u0026rsquo; personal information, business-sensitive data, and system operational data, among others. Privacy leakage refers to privacy information being obtained, used, or disclosed without authorization, including direct leakage via theft of raw data and indirect leakage inferred through correlation analysis.\u003c/p\u003e \u003cp\u003eCryptography is the theoretical foundation of privacy protection, providing the mathematical principles for encryption protection in distributed systems. It mainly includes three theories: symmetric encryption (e.g., AES, RC4), where encryption and decryption share a key\u0026mdash;high efficiency and low resource consumption, suitable for large-scale data transmission and storage, but key distribution and management are difficult in multi-node distributed environments and leakage risk is high; asymmetric encryption (e.g., RSA, ECC), which uses public-key encryption and private-key decryption\u0026mdash;simpler key management without a secure channel, but low efficiency and unsuitable for real-time processing of massive data; homomorphic encryption supports computation directly on encrypted data, enabling \u0026ldquo;data usable but not visible,\u0026rdquo; serving as a core enabler for privacy computing, but with high computational complexity and resource consumption and is currently difficult to adapt to large-scale real-time scenarios.\u003c/p\u003e \u003cp\u003ePrivacy computing theory concerns balancing privacy protection and value extraction. Key technologies include: federated learning, which enables multi-node joint training by sharing model parameters and includes horizontal, vertical, and transfer variants\u0026mdash;no raw data transmission and suitable for modeling with large numbers of nodes, but with high communication overhead and slow convergence; secure multi-party computation (SMPC), which uses secret sharing and related techniques to enable joint computation without data disclosure\u0026mdash;suitable for small-scale collaborative decision-making such as financial reconciliation, but with high computational complexity and poor real-time performance; differential privacy, which protects individual privacy by adding noise and includes centralized and local variants\u0026mdash;simple and low resource consumption, suitable for ultra-large-scale data collection, but noise affects data usability.\u003c/p\u003e \u003cp\u003eA comparison of theoretical characteristics of privacy computing technologies is shown in\u003c/p\u003e \u003cp\u003e(Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e\u0026ndash;3 Detailed Comparison of Theoretical Characteristics of Privacy Computing Technologies\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Taba\" border=\"1\"\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTechnology Type\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrivacy Protection Strength\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eComputational Efficiency\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eCommunication Overhead\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eData Usability\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eApplicable Node Scale\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eTypical Application Scenario\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFederated learning\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHigh\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMedium\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMedium\u0026ndash;high\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eHigh\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eLarge-scale\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eDistributed data modeling, e.g., collaborative medical data analysis\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSecure multi-party computation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eVery high\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLow\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eHigh\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eVery high\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eSmall-scale\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eDistributed collaborative decision-making, e.g., financial transaction reconciliation\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDifferential privacy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMedium\u0026ndash;high\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eVery high\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eVery low\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eMedium\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eUltra-large-scale\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eDistributed data collection, e.g., mobile crowdsensing\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHomomorphic encryption\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eVery high\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eVery low\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eVery high\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eVery high\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eMedium-scale\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eDistributed data querying, e.g., privacy information retrieval\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eA privacy-protection evaluation indicator system is used to assess effectiveness and mainly includes four categories of indicators: privacy protection strength measured by k-anonymity, l-diversity, t-closeness, and differential privacy budget ε; data usability evaluated by accuracy, completeness, and timeliness; system performance overhead including computation, communication, and storage overhead; and adaptability focusing on the ability to adapt to node heterogeneity and dynamic scaling. Indicator weights should be determined according to the scenario, using AHP or the entropy-weight method to achieve comprehensive evaluation.\u003c/p\u003e \u003c/div\u003e"},{"header":"3 Analysis of Privacy Leakage Problems and Causes in Large-Scale Distributed Systems","content":"\u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Manifestations of Privacy Leakage\u003c/h2\u003e \u003cp\u003eAcross the full lifecycle of data processing in large-scale distributed systems, privacy leakage exhibits multi-dimensional and complex characteristics, as detailed below:\u003c/p\u003e \u003cdiv id=\"Sec15\" class=\"Section3\"\u003e \u003ch2\u003e3.1.1 Privacy Leakage in the Data Collection Stage\u003c/h2\u003e \u003cp\u003eData collection is the first line of defense for privacy protection and also a high-risk stage. Core problems include: illegal collection\u0026mdash;distributed nodes may be maliciously controlled to steal sensitive information such as users\u0026rsquo; habits and locations; over-collection\u0026mdash;some applications collect information beyond business necessity to pursue data completeness (e.g., an energy monitoring system additionally collecting user identity data); identifier leakage\u0026mdash;collected data is not anonymized and contains unique identifiers such as ID numbers and device serial numbers, which can lead to precise privacy harms once leaked.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section3\"\u003e \u003ch2\u003e3.1.2 Privacy Leakage in the Data Transmission Stage\u003c/h2\u003e \u003cp\u003eData in distributed systems must be transmitted across nodes. The openness and multi-path nature of links increases leakage risks. Transmission eavesdropping refers to illegally intercepting data on the network; energy trading data and network situational awareness data may be stolen, causing commercial secrecy and cybersecurity issues. Data tampering refers to malicious modification of data during transmission, which can both distort data and indirectly leak sensitive information. Man-in-the-middle attacks occur when an attacker impersonates a legitimate node to intercept and forward data, stealing collaborative communications between nodes.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section3\"\u003e \u003ch2\u003e3.1.3 Privacy Leakage in the Data Processing Stage\u003c/h2\u003e \u003cp\u003eDuring multi-node collaborative computation, leakage risks arise in three main ways: malicious-node leakage\u0026mdash;edge nodes and fusion nodes may illegally store or leak sensitive data during processing; correlation-analysis leakage\u0026mdash;unauthorized privacy information can be derived from multi-source collaborative processing, for example, energy consumption data combined with location data to form user behavior trajectories; model-parameter leakage\u0026mdash;in privacy computing scenarios such as federated learning, attackers may infer raw training data by analyzing shared parameters.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section3\"\u003e \u003ch2\u003e3.1.4 Privacy Leakage in the Data Storage Stage\u003c/h2\u003e \u003cp\u003eThe multi-node nature of distributed storage increases risk points, including: storage intrusion\u0026mdash;distributed data centers and edge nodes may be attacked, leading to theft of large amounts of private data; data remanence\u0026mdash;data is not thoroughly erased after deletion and redundant backups are not synchronously deleted, enabling illegal recovery of residual data; access-control failure\u0026mdash;disordered permission management at storage nodes allows unauthorized users to access sensitive resources such as medical data and user information.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Analysis of Causes of Privacy Leakage\u003c/h2\u003e \u003cp\u003ePrivacy leakage in large-scale distributed systems results from the combined effects of system architecture, data characteristics, technical mechanisms, and management mechanisms. The main causes are as follows:\u003c/p\u003e \u003cdiv id=\"Sec20\" class=\"Section3\"\u003e \u003ch2\u003e3.2.1 Causes Related to System Architecture Characteristics\u003c/h2\u003e \u003cp\u003eLack of decentralized control: distributed systems have no unified privacy control center, and different nodes adopt different protection strategies; there is no unified protocol when constructing architectures such as peer-to-peer networks, and malicious nodes can exploit vulnerabilities to obtain data. Compatibility issues due to node heterogeneity: differences in hardware, software, and operating systems prevent unified deployment of privacy protection technologies; resource-constrained nodes cannot implement effective defenses. Blurred boundaries under dynamic scaling: when node scale changes dynamically, the privacy status of new nodes cannot be rapidly verified, and data-cleaning mechanisms for departing nodes are incomplete, leading to residual data leakage.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section3\"\u003e \u003ch2\u003e3.2.2 Causes Related to Data Characteristics\u003c/h2\u003e \u003cp\u003eFirst, multi-source heterogeneity increases protection difficulty: data includes structured, semi-structured, and unstructured types, with significant differences in privacy definition and protection methods, making unified policies difficult. Second, data correlations lead to indirect leakage: complex relationships among multi-source data mean that even if one source does not directly leak privacy, sensitive information can still be inferred through correlation analysis. Third, massive scale increases protection costs: PB-level or even EB-level data requires substantial computation, storage, and communication resources for privacy protection, and some systems reduce protection strength to control costs.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec22\" class=\"Section3\"\u003e \u003ch2\u003e3.2.3 Causes Related to Technical Mechanism Defects\u003c/h2\u003e \u003cp\u003eFirst, privacy protection technologies themselves have inherent limitations; second, data processing and privacy protection are decoupled; third, there are deficiencies in security protocols and encryption implementations.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec23\" class=\"Section3\"\u003e \u003ch2\u003e3.2.4 Causes Related to Missing Management Mechanisms\u003c/h2\u003e \u003cp\u003eFirst, standards and specifications are not unified: there is a lack of unified privacy protection standards for large-scale distributed systems, and implementation methods differ across scenarios, providing limited guidance for system design and operation. Second, node trust management is inadequate: lacking effective evaluation mechanisms to identify malicious nodes allows them to participate in data processing and sharing, stealing private information. Third, risk prevention and control mechanisms are missing: lacking routine privacy risk assessments means vulnerabilities cannot be discovered in time; after leakage events, emergency responses lag and cannot quickly contain risk spread.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"4 Theoretical Construction of a Privacy-Aware Data Processing Framework","content":"\u003cdiv id=\"Sec25\" class=\"Section2\"\u003e \u003ch2\u003e4.1 Framework Design Principles\u003c/h2\u003e \u003cp\u003eBased on the characteristics of large-scale distributed systems and privacy protection requirements, the framework follows four core principles:\u003c/p\u003e \u003cp\u003ePrivacy-first principle: privacy protection is embedded throughout the full data lifecycle. In the design of each layer and collaboration mechanisms, privacy security is prioritized, and methods such as source protection and process control are adopted to prevent illegal acquisition and leakage of private information.\u003c/p\u003e \u003cp\u003eCollaborative optimization principle: strike a balance among privacy protection, data processing efficiency, and usability. Protection strategies and processing parameters are dynamically adjusted, selecting appropriate technical solutions based on data sensitivity and node resource status to avoid imbalance among objectives.\u003c/p\u003e \u003cp\u003eScalability principle: adopt modular and layered design with standardized interfaces to support dynamic scaling of node count, data volume, and application scenarios. When adding nodes or scenarios, only corresponding modules or interface adaptations are needed without reconstructing the framework.\u003c/p\u003e \u003cp\u003eCompatibility principle: compatible with mainstream distributed architectures, data processing models, and privacy protection technologies to reduce system upgrade and migration costs.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec26\" class=\"Section2\"\u003e \u003ch2\u003e4.2 Layered Theoretical Architecture of the Framework\u003c/h2\u003e \u003cp\u003eFollowing these principles, a layered theoretical structure is constructed in which the Perception Layer, Processing Layer, Privacy Protection (Assurance) Layer, and Decision Layer collaborate top-down, deeply integrating the perception layer with data processing. The theoretical architecture is shown in Fig.\u0026nbsp;4\u0026thinsp;\u0026minus;\u0026thinsp;1.\u003c/p\u003e \u003cp\u003eFigure 4\u0026thinsp;\u0026minus;\u0026thinsp;1 Theoretical Architecture Diagram of the Privacy-Aware Data Processing Framework\u003c/p\u003e \u003cp\u003eThe core theoretical functions, key technological support, and design objectives of each layer are shown in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e4\u003c/span\u003e\u0026thinsp;\u0026minus;\u0026thinsp;1.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003e\u0026thinsp;\u0026minus;\u0026thinsp;1 Core Theoretical Functions and Design Objectives of Each Layer\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLayer\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCore Functions\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eKey Technology Support\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eDesign Objectives\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eCollaboration Relationship\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePerception layer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrivacy-aware collection, preprocessing, encrypted transmission\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLocal differential privacy, lightweight encryption, etc.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eControl leakage at the source; ensure secure collection and transmission\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eOutput data to the processing layer; receive instructions from the decision layer\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eProcessing layer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDistributed processing and fusion of private data\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFederated learning, secure multi-party computation, etc.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eEfficient processing; balance security and usability\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eReceive data from perception layer; output risks to assurance layer; receive decision-layer instructions\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePrivacy assurance layer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRisk assessment, strategy adaptation, security auditing\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRisk assessment system, dynamic adaptation, etc.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eReal-time monitoring; optimized protection\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eReceive risks from processing layer; advise decision layer; issue strategies to lower layers\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDecision layer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eOptimization scheduling, global strategy, monitoring\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMulti-objective optimization, resource scheduling, etc.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eGlobal optimized operation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eReceive reports from assurance layer; issue instructions to all layers; monitor globally\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec27\" class=\"Section2\"\u003e \u003ch2\u003e4.3 Core Collaborative Mechanisms of the Framework\u003c/h2\u003e \u003cp\u003eEfficient operation of the framework relies on three types of collaborative mechanisms to ensure deep integration among layers, nodes, and technical components:\u003c/p\u003e \u003cp\u003eBidirectional collaboration via standardized interfaces and message-passing protocols: top-down, the decision layer formulates global strategies and issues them to the privacy assurance layer, which decomposes them into specific strategies delivered to the perception and processing layers; bottom-up, the perception and processing layers report risk data to the privacy assurance layer, which aggregates and analyzes them to form reports and recommendations for the decision layer. Key supporting technologies include RESTful APIs, gRPC, message queues such as Kafka, and distributed consensus protocols such as Paxos and Raft to ensure real-time performance and reliability.\u003c/p\u003e \u003cp\u003ePrivacy-aware collaboration for multi-node characteristics: node trust evaluation is based on historical behavior data, using trust models to score nodes and assign permissions to exclude malicious nodes; node roles are dynamically assigned according to resource status and trust\u0026mdash;nodes with sufficient resources and high trust undertake core processing and fusion tasks, while resource-constrained nodes perform edge collection and processing; node data sharing adopts privacy computing theories to enable collaborative computation and secure fusion without leaking raw data.\u003c/p\u003e \u003cp\u003eCollaboration between privacy protection and data processing technologies: select different privacy protection technologies according to data sensitivity and application scenarios\u0026mdash;high-sensitivity data uses homomorphic encryption\u0026thinsp;+\u0026thinsp;federated learning, medium-sensitivity data uses differential privacy\u0026thinsp;+\u0026thinsp;symmetric encryption, and low-sensitivity data uses anonymization\u0026thinsp;+\u0026thinsp;access control. Privacy protection technologies are embedded into each stage of data processing to achieve deep integration. Based on risk assessment results and global optimization instructions, technology combinations are dynamically adjusted to balance privacy security, processing efficiency, and resource consumption.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec28\" class=\"Section2\"\u003e \u003ch2\u003e4.4 Analysis of Theoretical Advantages of the Framework\u003c/h2\u003e \u003cp\u003eCompared with existing frameworks, this framework has four theoretical advantages:\u003c/p\u003e \u003cp\u003eFull-lifecycle privacy awareness: privacy protection covers the entire process of collection, transmission, processing, and storage. Four-layer collaboration provides comprehensive protection, addressing the fragmentation of privacy protection in existing frameworks.\u003c/p\u003e \u003cp\u003eMulti-objective collaborative optimization: via dynamic matching mechanisms between the decision-layer optimization module and each layer, it coordinates privacy protection, processing speed, resource consumption, and data usability, breaking constraints of single-objective optimization.\u003c/p\u003e \u003cp\u003eHigh scalability and strong compatibility: modular, layered, and standardized-interface design enables dynamic expansion and compatibility with mainstream architectures, models, and technologies, reducing upgrade and migration costs.\u003c/p\u003e \u003cp\u003eDynamic adaptability: based on risk evaluation and technology collaboration mechanisms, strategies and technology combinations are adjusted according to system state, data characteristics, and privacy risks, adapting to the dynamicity and heterogeneity of distributed systems.\u003c/p\u003e \u003c/div\u003e"},{"header":"5 Analysis of Core Theories and Key Technologies of the Framework","content":"\u003cdiv id=\"Sec30\" class=\"Section2\"\u003e \u003ch2\u003e5.1 Core Theories and Technologies of the Perception Layer\u003c/h2\u003e \u003cdiv id=\"Sec31\" class=\"Section3\"\u003e \u003ch2\u003e5.1.1 Privacy-Aware Data Collection Theory\u003c/h2\u003e \u003cp\u003ePrivacy-aware data collection theory takes the principle of minimum necessary collection as the core and combines localized privacy protection technologies to address illegal collection, over-collection, and identifier leakage. It mainly includes three modules:\u003c/p\u003e \u003cp\u003eDynamic collection strategy optimization: develop different strategies based on data sensitivity (very high, high, medium, low). Use reinforcement learning to adjust collection scope, frequency, and granularity to balance privacy protection, usability, and collection cost.\u003c/p\u003e \u003cp\u003eCollected-data anonymization: combine k-anonymity, l-diversity, and t-closeness theories to jointly defend against identity and attribute inference attacks; k, l, and t values are adaptively adjusted based on data distribution.\u003c/p\u003e \u003cp\u003eLocal differential privacy (LDP): add noise locally before collection, with the core formula satisfying privacy budget ε constraints. For numerical, categorical, and high-dimensional data, respectively use the Laplace mechanism, exponential mechanism, and sparse vector techniques to balance privacy and usability.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003e\u0026thinsp;\u0026minus;\u0026thinsp;1 Comparison of Theoretical Characteristics of Perception-Layer Data Collection Technologies\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTheory/Technique\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrivacy Protection Strength\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eData Usability\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eComputational Overhead\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eApplicable Data Type\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eApplication Scenario\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDynamic collection strategy optimization\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;5%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eLow\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAll types\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eLarge-scale distributed multi-scenario data collection\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ek-anonymity (k\u0026thinsp;=\u0026thinsp;5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMedium (ε\u0026thinsp;=\u0026thinsp;2\u0026ndash;4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5%\u0026ndash;10%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMedium\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eStructured data\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eDisease surveillance data collection\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003el-diversity (l\u0026thinsp;=\u0026thinsp;3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMedium\u0026ndash;high (ε\u0026thinsp;=\u0026thinsp;1\u0026ndash;2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e10%\u0026ndash;15%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMedium\u0026ndash;high\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eStructured data\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eMedical data collection\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003et-closeness (t\u0026thinsp;=\u0026thinsp;0.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHigh (ε\u0026thinsp;=\u0026thinsp;0.5\u0026ndash;1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e15%\u0026ndash;20%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eHigh\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eStructured data\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eHighly sensitive business data collection\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLocal differential privacy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAdjustable (ε\u0026thinsp;=\u0026thinsp;0.1\u0026ndash;5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e8%\u0026ndash;25%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eLow\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eNumerical data\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eEnergy consumption data collection\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLocal differential privacy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAdjustable (ε\u0026thinsp;=\u0026thinsp;0.1\u0026ndash;5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e10%\u0026ndash;30%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eLow\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eCategorical data\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eUser behavior data collection\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec32\" class=\"Section3\"\u003e \u003ch2\u003e5.1.2 Data Transmission Security Theory\u003c/h2\u003e \u003cp\u003eBuild a triple mechanism of encrypted transmission, identity authentication, and integrity verification to adapt to node resource constraints. Key technologies include:\u003c/p\u003e \u003cp\u003eLightweight encrypted transmission: use improved AES (AES-CCM/GCM) and ECC algorithms, combining compression with encryption to reduce resource consumption.\u003c/p\u003e \u003cp\u003eTransmission-link identity verification: use distributed PKI digital certificates for identification and lightweight mutual authentication protocols to complete node-to-node verification and prevent man-in-the-middle attacks.\u003c/p\u003e \u003cp\u003eData integrity verification: use SM3/SHA-3 hash checks and ECC digital signatures; for highly sensitive data, use a double-hash mechanism to improve reliability.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec33\" class=\"Section3\"\u003e \u003ch2\u003e5.1.3 Privacy Risk Prevention and Control Theory of the Perception Layer\u003c/h2\u003e \u003cp\u003eBuild a multi-level risk prevention and control system:\u003c/p\u003e \u003cp\u003eRisk identification: extract collection, transmission, and node features; use lightweight machine-learning models to identify risks locally, categorizing severity into four levels: critical, high, medium, and low.\u003c/p\u003e \u003cp\u003eReal-time monitoring: coordinate local monitoring modules with distributed monitoring nodes to dynamically adjust the transmission frequency of monitoring data.\u003c/p\u003e \u003cp\u003eAnomaly response: formulate different strategies based on risk levels\u0026mdash;critical risks immediately stop collection/transmission; high risks switch to backup links; medium/low risks dynamically adjust protection strength.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec34\" class=\"Section2\"\u003e \u003ch2\u003e5.2 Core Theories and Technologies of the Processing Layer\u003c/h2\u003e \u003cp\u003eThe processing layer takes \"privacy computing as the core and distributed collaboration as the support\" and includes four core theoretical modules:\u003c/p\u003e \u003cdiv id=\"Sec35\" class=\"Section3\"\u003e \u003ch2\u003e5.2.1 Privacy-Preserving Data Preprocessing Theory\u003c/h2\u003e \u003cp\u003eEmbed privacy protection technologies into cleaning, transformation, and integration processes:\u003c/p\u003e \u003cp\u003ePrivacy-aware cleaning: use federated learning to impute missing values, distributed anomaly detection, and encrypted-hash deduplication; raw data remains locally stored.\u003c/p\u003e \u003cp\u003ePrivacy-aware transformation: perform dynamic desensitization based on sensitivity levels; use federated learning for feature extraction and distributed PCA for dimensionality reduction.\u003c/p\u003e \u003cp\u003ePrivacy-aware integration: adopt decentralized architecture\u0026thinsp;+\u0026thinsp;ABE encryption; use secure multi-party computation to resolve data conflicts.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003e\u0026thinsp;\u0026minus;\u0026thinsp;2 Comparison of Theoretical Characteristics of Processing-Layer Data Preprocessing Technologies\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTechnology Type\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrivacy Protection Method\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eData Usability\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eComputational Overhead\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eCommunication Overhead\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFederated-learning imputation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eModel sharing\u0026thinsp;+\u0026thinsp;local computation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eHigh (\u0026gt;\u0026thinsp;95%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMedium\u0026ndash;high\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eMedium\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDistributed anomaly detection\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLocal detection\u0026thinsp;+\u0026thinsp;identifier sharing\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eHigh (\u0026gt;\u0026thinsp;90%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMedium\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eLow\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEncrypted-hash deduplication\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHash comparison\u0026thinsp;+\u0026thinsp;local storage\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eVery high (100%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eLow\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eLow\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDynamic data desensitization\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSubstitution/generalization/masking\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMedium\u0026ndash;high (85%\u0026ndash;95%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eLow\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eNone\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFederated feature extraction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLocal extraction\u0026thinsp;+\u0026thinsp;encrypted transmission\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMedium (80%\u0026ndash;90%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMedium\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eMedium\u0026ndash;high\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eABE-encrypted integration\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAttribute-based encryption\u0026thinsp;+\u0026thinsp;distributed storage\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eHigh (\u0026gt;\u0026thinsp;90%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMedium\u0026ndash;high\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eMedium\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSMPC-based conflict resolution\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eJoint computation\u0026thinsp;+\u0026thinsp;noise addition\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMedium (80%\u0026ndash;85%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eHigh\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eMedium\u0026ndash;high\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec36\" class=\"Section3\"\u003e \u003ch2\u003e5.2.2 Distributed Privacy Computing Theory\u003c/h2\u003e \u003cp\u003eAchieve \"data usable but not visible\" through synergy among three technologies:\u003c/p\u003e \u003cp\u003eFederated learning: adapt to layered, peer-to-peer, and microservices architectures; enhance privacy via parameter perturbation, model compression, and secure aggregation; use adaptive training strategies for heterogeneous nodes.\u003c/p\u003e \u003cp\u003eSecure multi-party computation: split complex tasks and data into shares; improve efficiency via offline preprocessing\u0026thinsp;+\u0026thinsp;hybrid protocols; support dynamic node join/leave.\u003c/p\u003e \u003cp\u003eHomomorphic encryption collaboration: choose PHE, LHE, or FHE algorithms according to the computation task; schedule ciphertext computation in a distributed manner and perform distributed decryption to avoid private-key leakage.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec37\" class=\"Section3\"\u003e \u003ch2\u003e5.2.3 Secure Data Fusion Theory\u003c/h2\u003e \u003cp\u003eBuild a decentralized fusion architecture; key technologies include:\u003c/p\u003e \u003cp\u003eDistributed fusion architecture: three-level fusion (edge\u0026ndash;regional\u0026ndash;global) combined with peer-node fusion, dynamically selecting core fusion nodes.\u003c/p\u003e \u003cp\u003ePrivacy-preserving fusion algorithms: encrypted-data fusion, federated-learning fusion, distributed compressive-sensing fusion, and differential-privacy fusion.\u003c/p\u003e \u003cp\u003eFusion-result protection: ABE access control, encrypted storage, and dynamically desensitized graded sharing.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec38\" class=\"Section3\"\u003e \u003ch2\u003e5.2.4 Privacy-Preserving Storage Theory\u003c/h2\u003e \u003cp\u003eBuild a triple system of \"encrypted storage\u0026thinsp;+\u0026thinsp;access control\u0026thinsp;+\u0026thinsp;data sanitization\":\u003c/p\u003e \u003cp\u003eDistributed encrypted storage: data sharding encryption\u0026thinsp;+\u0026thinsp;hybrid encryption; blockchain-assisted storage indexes and logs.\u003c/p\u003e \u003cp\u003eFine-grained access control: integrate RBAC, ABE, and zero-trust mechanisms to achieve precise permission control.\u003c/p\u003e \u003cp\u003eSecure data sanitization: distributed collaborative sanitization; overwriting/shredding physical media; offline archiving for highly sensitive data.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec39\" class=\"Section2\"\u003e \u003ch2\u003e5.3 Core Theories and Technologies of the Privacy Assurance Layer\u003c/h2\u003e \u003cp\u003eThe privacy assurance layer serves as the framework\u0026rsquo;s \"privacy security hub.\" Its main purpose is to dynamically control privacy risks across the full data processing lifecycle. With the principles of \"risk-driven, dynamic adaptation, end-to-end auditing, and rapid response,\" it includes four parts: privacy risk assessment, privacy protection strategy adaptation, collaborative scheduling of multiple privacy technologies, security auditing, and emergency response, ensuring that privacy protection effectiveness consistently meets requirements.\u003c/p\u003e \u003cdiv id=\"Sec40\" class=\"Section3\"\u003e \u003ch2\u003e5.3.1 Privacy Risk Assessment Theory\u003c/h2\u003e \u003cp\u003ePrivacy risk assessment is the foundation of the privacy assurance layer. By building a scientific indicator system and assessment model, it quantitatively analyzes full-lifecycle privacy risks, providing a basis for strategy adaptation and emergency response.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e\n\u003ch3\u003e1) Full-lifecycle risk assessment indicator system\u003c/h3\u003e\n\u003cp\u003eAn indicator system is constructed across the four stages\u0026mdash;data collection, transmission, processing, and storage\u0026mdash;covering three dimensions: risk sources, risk impacts, and risk controls. Indicator weights are calculated using the entropy-weight method; the processing stage involves multi-node collaboration and data fusion, has the most complex risks, and thus receives the highest weight. Specific indicators are shown in Table\u0026nbsp;\u003cspan refid=\"Tab7\" class=\"InternalRef\"\u003e5\u003c/span\u003e\u0026thinsp;\u0026minus;\u0026thinsp;3.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab6\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003e\u0026thinsp;\u0026minus;\u0026thinsp;3 Full-Lifecycle Privacy Risk Assessment Indicator System\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAssessment Stage\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRisk Source Dimension\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRisk Impact Dimension\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eRisk Control Dimension\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eIndicator Weight (Entropy-weight method)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eData collection\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eProbability of illegal collection; degree of over-collection; probability of identifier leakage\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eData sensitivity; scope of leakage impact; compliance risk level\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eRationality of collection strategy; anonymization strength; LDP noise strength\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.22\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eData transmission\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eProbability of link eavesdropping; probability of data tampering; probability of MITM attacks\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eSensitivity of transmitted data; leakage propagation speed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eEncryption strength; authentication effectiveness; integrity-check success rate\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.25\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eData processing\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eProportion of malicious nodes; probability of correlation-inference leakage; probability of model-parameter leakage\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eSensitivity of processed data; impact of fusion-result leakage\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSecurity of privacy computing technologies; trust level of processing nodes\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.28\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eData storage\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eProbability of node intrusion; probability of data remanence; probability of access-control failure\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eSensitivity of stored data; data storage period\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eEncrypted storage strength; access-control granularity; completeness of data sanitization\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.25\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e\n\u003ch3\u003e2) Risk assessment model theory\u003c/h3\u003e\n\u003cp\u003eA multi-level, multi-method fusion model is used for accurate quantification: a hierarchical assessment model with target, criterion, and indicator layers uses AHP and fuzzy comprehensive evaluation to quantify risk levels; machine-learning risk prediction models train random forests, SVM, LSTM, etc. on historical and real-time data, with federated learning improving generalization; a dynamic risk assessment model uses a sliding-window mechanism to update data in real time, adjust indicator weights and risk thresholds, and ensure timeliness.\u003c/p\u003e\n\u003ch3\u003e3) Risk level determination theory\u003c/h3\u003e\n\u003cp\u003eBased on model outputs, a threshold method classifies overall privacy risk into four levels. Each level corresponds to specific risk descriptions and handling requirements, as shown in Table\u0026nbsp;\u003cspan refid=\"Tab7\" class=\"InternalRef\"\u003e5\u003c/span\u003e\u0026thinsp;\u0026minus;\u0026thinsp;4.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab7\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003e\u0026thinsp;\u0026minus;\u0026thinsp;4 Privacy Risk Level Determination Standards\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRisk Level\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eQuantified Score\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRisk Description\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eHandling Requirements\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCritical risk\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e80\u0026ndash;100\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eSevere leakage hazard; may cause large-scale sensitive data leakage and non-compliance\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eImmediate emergency response; suspend business and investigate\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHigh risk\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e60\u0026ndash;79\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eObvious leakage risk; may cause partial sensitive data leakage and impact operations\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eEnable high-level protection; investigate within a time limit\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMedium risk\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e40\u0026ndash;59\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePotential leakage risk; minor impact; not involving core business\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eEnable medium-level protection; continuous monitoring and optimization\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLow risk\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0\u0026ndash;39\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLow and controllable risk; no impact on operations and data security\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMaintain basic protection; periodic assessment\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cdiv id=\"Sec44\" class=\"Section3\"\u003e \u003cdiv class=\"Heading\"\u003e5.3.2 Privacy Protection Strategy Adaptation Theory\u003c/div\u003e \u003cp\u003eBased on risk assessment results, privacy protection strategy adaptation dynamically adjusts privacy strategies and technology combinations to achieve a precise \"risk\u0026ndash;strategy\" match, addressing the static nature and weak targeting of traditional strategies.\u003c/p\u003e \u003cp\u003eStrategy adaptation decision theory\u003c/p\u003e \u003cp\u003eWith \"meeting privacy protection strength requirements, minimizing system performance overhead, and maximizing data usability\" as a multi-objective optimization direction, the objective function is as follows:\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWhere C is system performance overhead (computation\u0026thinsp;+\u0026thinsp;communication\u0026thinsp;+\u0026thinsp;storage), A is data usability, S is actual privacy protection strength, S_min is the minimum protection strength required by risk, and w is a weight coefficient adjusted according to business needs.\u003c/p\u003e \u003cp\u003eDecision algorithms: use genetic algorithms or particle swarm optimization to search for the optimal combination of privacy protection strategies (e.g., selection of encryption techniques, selection of privacy computing techniques, parameter settings). For example, when the risk level is high, the optimization prioritizes meeting S_min and selects combinations with stronger privacy protection (e.g., \"homomorphic encryption\u0026thinsp;+\u0026thinsp;federated learning\"); when the risk level is low, the optimization prioritizes reducing C and improving A, selecting lightweight combinations (e.g., \"symmetric encryption\u0026thinsp;+\u0026thinsp;anonymization\").\u003c/p\u003e \u003cp\u003eDynamic strategy adjustment theory\u003c/p\u003e \u003cp\u003eSupport adaptive adjustments under three scenarios: risk-driven adjustment (increase/decrease protection strength as risk level rises/falls); system-state adaptation (switch to lightweight techniques when resource utilization is too high; optimize collaboration strategies as node count increases); business-demand adaptation (dynamically balance privacy and usability according to data accuracy or compliance requirements).\u003c/p\u003e \u003cp\u003eStrategy issuance and execution theory\u003c/p\u003e \u003cp\u003eUse standardized policy description languages (XACML, JSON) to translate policies into executable commands; issue policies through a three-level distributed mechanism of \"decision layer \u0026ndash; privacy assurance layer \u0026ndash; perception/processing layer,\" combined with sharded issuance to reduce communication overhead [10][12]; establish a policy execution verification mechanism\u0026mdash;after nodes feedback execution results, the privacy assurance layer verifies compliance, and abnormal nodes will be restricted from participating in data processing.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec45\" class=\"Section3\"\u003e \u003cdiv class=\"Heading\"\u003e5.3.3 Collaborative Scheduling Theory for Multiple Privacy Technologies\u003c/div\u003e \u003cp\u003eThe core is to break through the limitations of single technologies through combination optimization and dynamic adaptation, achieving a three-way balance. Key contents are as follows:\u003c/p\u003e \u003cp\u003eCore combination modes: build three streamlined modes by sensitivity level, covering the entire link and supporting recombination:\u003c/p\u003e \u003cp\u003eHigh-strength protection (fully homomorphic encryption\u0026thinsp;+\u0026thinsp;secure multi-party computation\u0026thinsp;+\u0026thinsp;sharded encrypted storage): for extremely sensitive data, ε\u0026thinsp;\u0026le;\u0026thinsp;1, usability 70%\u0026ndash;80%.\u003c/p\u003e \u003cp\u003eBalanced protection (federated learning\u0026thinsp;+\u0026thinsp;partially homomorphic encryption\u0026thinsp;+\u0026thinsp;ABE access control): for medium\u0026ndash;high sensitive data, ε\u0026thinsp;=\u0026thinsp;1\u0026ndash;2, usability 80%\u0026ndash;90%.\u003c/p\u003e \u003cp\u003eLightweight protection (k-anonymity with k\u0026thinsp;\u0026ge;\u0026thinsp;5\u0026thinsp;+\u0026thinsp;AES encryption\u0026thinsp;+\u0026thinsp;RBAC access control): for low\u0026ndash;medium sensitive data, usability 90%\u0026ndash;95%.\u003c/p\u003e \u003cp\u003eIntelligent scheduling mechanism: model with a DRL algorithm as an MDP. The state space includes risk level, resource utilization, and data sensitivity; the action space covers 3 standard combinations and 17 derived combinations. The reward function centers on privacy compliance. After 1,200 training rounds, it converges with a strategy selection accuracy of 92.3%, privacy compliance rate improved by 18%, and system overhead reduced by 22%.\u003c/p\u003e \u003cp\u003eOptimization constraints and triggers: the objective is to maximize privacy strength, minimize overhead, and maximize usability, with constraints including privacy compliance thresholds and resource limits. Scheduling triggers include changes in risk level, resource fluctuations, and business switching to ensure real-time adaptation.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec46\" class=\"Section2\"\u003e \u003ch2\u003e5.4 Core Theories and Technologies of the Decision Layer\u003c/h2\u003e \u003cp\u003eAs the global optimization hub, the decision layer uses a three-dimensional multi-objective optimization model of \"privacy\u0026ndash;efficiency\u0026ndash;usability\" to integrate operational status and risk data from all layers, formulating global resource scheduling, security strategies, and processing strategies. Based on distributed game theory to balance node interests, it dynamically allocates computation and storage resources, optimizes privacy-computing aggregation cycles and data transmission paths, breaks inter-layer collaboration barriers, and achieves globally optimal framework operation.\u003c/p\u003e \u003cp\u003eWhen scaling from 200 to 1000 nodes, compliance remained above 91% and overhead increase remained within 3%, demonstrating scalability.\u003c/p\u003e \u003c/div\u003e"},{"header":"6. Simulation and Experimental Evaluation","content":"\u003cdiv id=\"Sec48\" class=\"Section2\"\u003e \u003ch2\u003e6.1 Experimental Setup\u003c/h2\u003e \u003cp\u003eTo validate the effectiveness of the proposed privacy-aware framework, simulation experiments were conducted in a distributed computing environment.\u003c/p\u003e \u003cdiv id=\"Sec49\" class=\"Section3\"\u003e \u003ch2\u003e6.1.1 Simulation Environment\u003c/h2\u003e \u003cp\u003eThe simulation platform was implemented using Python 3.10 with PyTorch for DRL modeling. The experiments were executed on a workstation equipped with:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eCPU: 16-core Intel Xeon\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eMemory: 64 GB RAM\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eOperating System: Ubuntu 22.04\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eNetwork Simulation: Distributed node communication latency randomly sampled between 5\u0026ndash;50 ms\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec50\" class=\"Section3\"\u003e \u003ch2\u003e6.1.2 Distributed System Configuration\u003c/h2\u003e \u003cp\u003eThe simulated large-scale distributed system consisted of:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eNode scale: 200\u0026ndash;1000 heterogeneous nodes\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eData types: structured (40%), numerical time-series (35%), semi-structured logs (25%)\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eSensitivity levels: High (30%), Medium (45%), Low (25%)\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eRisk state updates: every 50 simulation steps\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eEach node was assigned dynamic resource constraints (CPU utilization 40\u0026ndash;85%) to simulate heterogeneous environments.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec51\" class=\"Section2\"\u003e \u003ch2\u003e6.2 Baseline Strategies\u003c/h2\u003e \u003cp\u003eTo evaluate the proposed framework, it was compared with two baseline approaches:\u003c/p\u003e \u003cp\u003eBaseline 1: Static Privacy Strategy\u003c/p\u003e \u003cp\u003eA fixed privacy protection strategy was deployed regardless of risk variation.\u003c/p\u003e \u003cp\u003eBaseline 2: Single-Technology Deployment\u003c/p\u003e \u003cp\u003eOnly one privacy mechanism (federated learning\u0026thinsp;+\u0026thinsp;symmetric encryption) was applied without adaptive scheduling.\u003c/p\u003e \u003cp\u003eProposed Method\u003c/p\u003e \u003cp\u003eThe proposed DRL-based multi-privacy collaborative scheduling mechanism dynamically selected optimal strategy combinations.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec52\" class=\"Section2\"\u003e \u003ch2\u003e6.3 Evaluation Metrics\u003c/h2\u003e \u003cp\u003eThree key performance indicators were used:\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003ePrivacy Compliance Rate (PCR) Percentage of operations satisfying the minimum privacy threshold S\u0026thinsp;\u0026ge;\u0026thinsp;S_min.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eSystem Overhead (SO) Normalized computational\u0026thinsp;+\u0026thinsp;communication overhead\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eData Usability (DU) Output data accuracy retention ratio after privacy protection\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec53\" class=\"Section2\"\u003e \u003ch2\u003e6.4 Experimental Results\u003c/h2\u003e \u003cdiv id=\"Sec54\" class=\"Section3\"\u003e \u003ch2\u003e6.4.1 Privacy Compliance\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Tabb\" border=\"1\"\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eStrategy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrivacy Compliance Rate\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eStatic Strategy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e74.6%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSingle Technology\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e81.2%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eProposed Framework\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e92.3%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"2\"\u003eThe proposed framework improved compliance by:\u003c/td\u003e\u003c/tr\u003e \u003ctr\u003e\u003ctd colspan=\"2\"\u003e● +\u0026thinsp;17.7% compared with Static Strategy\u003c/td\u003e\u003c/tr\u003e \u003ctr\u003e\u003ctd colspan=\"2\"\u003e● +\u0026thinsp;11.1% compared with Single Technology\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThis demonstrates the effectiveness of risk-driven adaptive privacy scheduling.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec55\" class=\"Section3\"\u003e \u003ch2\u003e6.4.2 System Overhead\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Tabc\" border=\"1\"\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eStrategy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNormalized Overhead\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eStatic Strategy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSingle Technology\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.91\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eProposed Framework\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.78\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe DRL-based adaptive scheduling reduced system overhead by approximately:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e22% compared with static deployment\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eThis is achieved by avoiding unnecessary high-intensity encryption when risk level is low.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec56\" class=\"Section3\"\u003e \u003ch2\u003e6.4.3 Data Usability\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Tabd\" border=\"1\"\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eStrategy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eData Usability\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eStatic Strategy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e82.4%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSingle Technology\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e85.1%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eProposed Framework\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e90.2%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe proposed method maintains higher usability due to dynamic parameter tuning of privacy mechanisms (e.g., adaptive ε selection in differential privacy).\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec57\" class=\"Section2\"\u003e \u003ch2\u003e6.5 DRL Convergence Analysis\u003c/h2\u003e \u003cp\u003eThe DRL scheduling agent was trained over 1200 episodes.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eConvergence observed after approximately 850 episodes\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eReward stabilization variance\u0026thinsp;\u0026lt;\u0026thinsp;3%\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eFinal strategy selection accuracy: 92.3%\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eThe reward function balanced:\u003c/p\u003e \u003cp\u003eR\u0026thinsp;=\u0026thinsp;αS - βC\u0026thinsp;+\u0026thinsp;γA\u003c/p\u003e \u003cp\u003ewhere S represents privacy strength,\u003c/p\u003e \u003cp\u003eC represents system overhead,\u003c/p\u003e \u003cp\u003eand A represents data usability.\u003c/p\u003e \u003cp\u003eThe model demonstrated stable policy selection under dynamic risk fluctuations.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec58\" class=\"Section2\"\u003e \u003ch2\u003e6.6 Scalability Analysis\u003c/h2\u003e \u003cp\u003eThe framework was tested under increasing node scale:\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Tabe\" border=\"1\"\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNodes\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCompliance\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eOverhead\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e200\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e93.1%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.75\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e500\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e92.6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.77\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e1000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e91.8%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.80\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003ePerformance degradation remained within 3%, confirming scalability suitability for large-scale distributed environments.\u003c/p\u003e \u003c/div\u003e"},{"header":"7. Discussion","content":"\u003cp\u003eThe simulation results indicate:\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eThe adaptive scheduling mechanism significantly improves privacy compliance.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eDynamic strategy selection reduces redundant computational costs.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eThe framework maintains stable performance under heterogeneous node scaling.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cp\u003eLimitations include:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eSynthetic dataset usage\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eLack of real-world deployment validation\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eSimplified adversarial model\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eFuture work will extend validation to real distributed energy and IoT datasets.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eFocusing on privacy leakage in large-scale distributed systems, this paper follows the logical chain of \"problem\u0026ndash;cause\u0026ndash;measure\" and constructs a complete theoretical system for a privacy-aware data processing framework. Key conclusions are as follows:\u003c/p\u003e \u003cp\u003eA four-layer collaborative architecture of \"perception\u0026ndash;processing\u0026ndash;privacy assurance\u0026ndash;decision\" is proposed, establishing a full-lifecycle privacy-aware theoretical system and forming a closed-loop paradigm of \"source protection\u0026ndash;process control\u0026ndash;global optimization,\" addressing the fragmentation of privacy protection and its decoupling from data processing in traditional research.\u003c/p\u003e \u003cp\u003eLayer-specific innovations are achieved: the perception layer constructs source protection theory, the processing layer forms process protection theory, the privacy assurance layer establishes dynamic control theory, and the decision layer proposes global optimization theory. The layers support each other to enable deep synergy between privacy and performance.\u003c/p\u003e \u003cp\u003eA collaborative scheduling theory for multiple privacy technologies is innovated by integrating encryption, privacy computing, and other core technologies. Through a DRL-based intelligent scheduling algorithm, it enables scenario-accurate adaptation, balancing privacy protection, system performance, and data usability, and overcoming the limitations of single technologies.\u003c/p\u003e \u003cp\u003eThe framework theory is compatible with existing architectures and processing models and can be applied to scenarios such as the energy Internet and smart homes. It provides unified theoretical guidance for practical systems and can effectively reduce leakage risks.\u003c/p\u003e \u003cp\u003eThis theoretical system fills a related theoretical gap. Future work can further explore collaborative optimization of technologies in dynamic heterogeneous environments and cross-domain privacy compliance adaptation to improve practicality and extensibility.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eFunding:\u003c/h2\u003e \u003cp\u003eNot applicable.\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eJ.T. conceived the study, designed the methodology, conducted the experiments, and wrote the main manuscript text. Z.L. assisted with data analysis and implementation. R.L. contributed to literature review and manuscript revision. All authors reviewed and approved the final manuscript.\u003c/p\u003e\u003ch2\u003eAcknowledgement\u003c/h2\u003e\u003cp\u003eThe authors would like to thank colleagues for valuable discussions.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe datasets generated and analyzed during the current study are available from the corresponding author on reasonable request.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eQue Hang, Wang Yizhuo, Jin Zhuohao, et al. A review of distributed ISAC data collection/processing and resource allocation research [J]. Mobile Communications, 2025, 49(12): 99–108.\u003c/li\u003e\n\u003cli\u003eHuang Y, Jiang J, Gao Z, et al. PIUL-DERs: Physics-informed pseudo-labeling unsupervised learning for smart homes with distributed energy resources [J]. Applied Energy, 2026, 403(PA): 127045. DOI:10.1016/J.APENERGY.2025.127045.\u003c/li\u003e\n\u003cli\u003eZhang R, Yang Y, Bie Z, et al. Distributed peer to peer transaction framework for heterogeneous energy trading in urban energy internet-integrated micro-energy grid systems with diverse preference aware [J]. International Journal of Electrical Power and Energy Systems, 2025, 173: 111332. DOI:10.1016/J.IJEPES.2025.111332.\u003c/li\u003e\n\u003cli\u003eLiu Y, Zhou D, Cheng L. Distributed online fusion learning of multi-source network situation awareness data [J]. Information Fusion, 2026, 127(PC): 103906. DOI:10.1016/J.INFFUS.2025.103906.\u003c/li\u003e\n\u003cli\u003eZhu Meiyi. Research on key technologies of wireless distributed inference systems [D]. Beijing University of Posts and Telecommunications, 2025.\u003c/li\u003e\n\u003cli\u003eNi Huiming. Design of a disease surveillance information system based on a distributed control system [J]. Wireless Internet Technology, 2025, 22(13): 39–43.\u003c/li\u003e\n\u003cli\u003eHe Lili, Jiang Sheng, Guan Xinru, et al. Privacy protection in mobile crowdsensing: explorations in secure sensing, development, and future [J]. Computer Applications Research, 2025, 42(11): 3201–3214. DOI:10.19734/j.issn.1001-3695.2025.04.0092.\u003c/li\u003e\n\u003cli\u003eCai Liuping, Chen Huihong. Design and implementation of a distributed environmental monitoring platform based on the Internet of Things [J]. Computer Programming Skills \u0026amp; Maintenance, 2025(06): 14–16 + 27.\u003c/li\u003e\n\u003cli\u003eWei Hao, Zhang Mengjie, Wang Dongming. Mobility management for distributed collaborative integrated sensing-and-communication networks [J]. Telecommunications Science, 2025, 41(03): 73–86.\u003c/li\u003e\n\u003cli\u003eWan Kaicheng. Research on parallel computing and distributed storage technologies in big data processing [J]. Information \u0026amp; Computer (Theory Edition), 2024, 36(19): 166–168.\u003c/li\u003e\n\u003cli\u003eTang Ronghua. Design and implementation of a campus network security situational awareness system based on big data [J]. Electronic Components and Information Technology, 2024, 8(09): 144–147.\u003c/li\u003e\n\u003cli\u003eShu Xinyue. Research on resource scheduling optimization for geographically distributed data centers [D]. Chongqing University, 2024.\u003c/li\u003e\n\u003cli\u003eYang Xiaolan. Construction of a distributed network massive data processing system based on cloud computing technology [J]. Wireless Internet Technology, 2023, 19(02): 68–70.\u003c/li\u003e\n\u003cli\u003eWang Ji. Research on intelligent optimization methods for information processing oriented to autonomous collaboration of unmanned swarms [D]. National University of Defense Technology, 2019.\u003c/li\u003e\n\u003cli\u003eKou Lan, Liu Ning, Huang Hongcheng, et al. A privacy-preserving data fusion algorithm based on distributed compressive sensing and hash functions [J]. Computer Applications Research, 2020, 37(01): 239–244.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"large-scale distributed systems, privacy awareness, data processing","lastPublishedDoi":"10.21203/rs.3.rs-8923292/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8923292/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe widespread adoption of large-scale distributed systems has promoted the efficient circulation and processing of multi-source data; however, the risk of data privacy leakage has become a core bottleneck constraining their development. This paper focuses on theoretical research on privacy-aware data processing. Based on the data processing characteristics of distributed systems and the core requirements of privacy protection, it constructs a privacy-aware data processing framework at the theoretical level, clarifies the framework\u0026rsquo;s core layers and theoretical connotations, and systematically analyzes the theoretical linkages among key enabling technologies. Through a literature review of existing related work, this paper identifies the current research status and gaps in distributed data processing and privacy protection, providing theoretical support for the collaborative optimization of data processing and privacy protection in large-scale distributed systems.\u003c/p\u003e","manuscriptTitle":"Research on a Privacy-Aware Data Processing Framework for Large-Scale Distributed Systems","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-06 12:09:20","doi":"10.21203/rs.3.rs-8923292/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-04-28T08:07:20+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-20T02:45:01+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"28640343497442951141454289974800152321","date":"2026-04-20T02:33:59+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"250614065995441816492913984312384209567","date":"2026-04-15T00:58:47+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-03-21T13:46:43+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"194946053650933522282681930012383865689","date":"2026-03-02T05:56:58+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-03-02T05:34:39+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-02-24T16:27:11+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-02-23T11:44:20+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-02-23T11:43:53+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2026-02-20T07:16:34+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"73f41683-2d91-4653-a9ef-436db207a7cd","owner":[],"postedDate":"March 6th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"in-revision","subjectAreas":[{"id":63887390,"name":"Physical sciences/Engineering"},{"id":63887391,"name":"Physical sciences/Mathematics and computing"}],"tags":[],"updatedAt":"2026-04-28T08:23:05+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-06 12:09:20","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8923292","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8923292","identity":"rs-8923292","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.