We present a comprehensive analysis of privacy attacks and countermeasures in data-driven systems. We systematically categorize attacks targeting three domains: anonymous data (linkage and structural attacks), statistical aggregates (reconstruction and differential attacks), and privacy-preserving models (extraction, reconstruction, membership inference, and inversion attacks). For each category, we analyze attack methodologies, adversary capabilities, and vulnerability mechanisms. We further evaluate countermeasures including perturbation techniques, randomization methods, query auditing, and model-level defenses, examining their effectiveness and inherent privacy-utility tradeoffs. Our analysis reveals that while differential privacy offers strong theoretical guarantees, it faces implementation challenges and potential vulnerabilities to emerging attacks. We identify critical research directions and provide researchers and practitioners with a structured framework for understanding privacy resilience in increasingly complex data ecosystems.
The K-line is one of the most widely recognized technical indicators, garnering significant attention from investors and serving as a prevalent reference point for stock market investments. This paper provides an innovative investment strategy rooted in a complex network that is shaped by the correlations among K-lines. Its monthly return reaches an impressive 5.4 % for constituents of CSI 300 index (one of the most popular China Securities Indexes), significantly outperforming the market. The analysis also reveals i) utilizing the K-lines network, a portfolio tracking the market can be effectively assembled with ten to twenty selected stocks and ii) portfolios constructed from low-centrality nodes surpass those constructed from high-centrality nodes. This paper provides a good solution to the costly management of portfolio construction.
Crowdsourcing labeling plays an indispensable role in modern data-driven applications, as it efficiently assists in tackling large-scale labeling tasks. To ensure the utility of data, recent works of crowdsourcing labeling deploy various quality control strategies. However, these strategies have exposed crowdsourcing workers to privacy threats. Current efforts on privacy protection in crowdsourcing mostly focus on addressing privacy concerns in spatial crowdsourcing rather than crowdsourcing labeling. Moreover, they overlook the issue of balancing privacy and data quality. In light of this, we propose a novel differential privacy crowdsourcing labeling algorithm with quality control to protect worker privacy. In this work, we incorporate workers’ confidence into the quality control strategy and provide a differentially private selection mechanism to select capable workers. This mechanism can strictly preserve privacy while maintaining sufficient data utility. To further ensure label quality, we design a consensus-based label aggregation algorithm that filters out unreliable answers, thus mitigating the impact of noise introduced by the privacy mechanism on data utility. The proposed algorithm reduces the dependence of the existing quality control strategy on historical information and realizes worker-level privacy protection through the randomness introduced by differential privacy. Experimental results on six real-world datasets demonstrate that our algorithm achieves a good balance between high-quality labels and strict privacy protection.
Users typically have varying privacy requirements when submitting their personal data to data collectors. To address this need, researchers have expanded existing Local Differential Privacy (LDP) methods into a personalized format, allowing users to select their privacy budget according to their individual privacy demands. However, these schemes uniformly apply an LDP method for privacy protection, neglecting the fact that different LDP methods yield varying data utility under the same privacy budget. To enhance data utility within personalized privacy protection, this paper proposes an adaptive utility optimization framework for personalized LDP and applies it to frequency estimation. Firstly, we present a mechanism for the adaptive matching of LDPs with minimal errors for users. This is accomplished by comparing the theoretical errors of different LDPs under personalized privacy budgets. Subsequently, we integrate this approach with a weighted combination optimization method to propose a novel adaptive utility optimization technique for personalized LDP. We also provide a theoretical proof of the effectiveness of our method in optimizing both privacy and utility. Finally, our method has been experimentally validated, demonstrating its efficacy as a tool for utility optimization, outperforming existing methods.
This study is to ascertain the implications of oil price shocks for Chinese green bond spreads. Green and non-green bonds of the same firm are matched in our analysis to clarify the link between oil shocks and green bonds is associated with their greenness. We find robust evidence that the green bond spreads widen compared to non-green bond spreads during oil demand shocks, whereas no significant performance difference is observed between green and non-green bonds during oil supply shocks. Further cross-sectional analysis reveals that the consequences of oil shocks on green bond yield spreads differ significantly among industries based on their oil use intensity and firms based on their greenness. Collectively, these findings offer new insights into the implications of oil shocks triggered by changes in demand and supply on green bond spreads.
Data releasing and sharing between several fields has became inevitable tendency in the context of big data. Unfortunately, this situation has clearly caused enormous exposure of sensitive and private information. Along with massive privacy breaches, privacy-preservation issues were brought into sharp focus and privacy concerns may prevent people from providing their personal data. To meet the requirements of privacy protection, such a problem has been extensively studied. However, privacy protection of sensitive information should not prevent data users from conducting valid analyses of the released data. We propose a novel algorithm in this paper, named Data Release under Adjustable Privacy-utility Equilibrium (DRAPE), to address this problem. We handle the privacy versus utility tradeoff in the data release problem by breaking sensitive associations among variables while maintaining the correlations of nonsensitive variables. Furthermore, we quantify the impact of the proposed privacy-preserving method in terms of correlation preservation and privacy level, and thereby develop an optimization model to fulfil data privacy and data utility constraints. The proposed approach is not only able to provide a better privacy levels control scheme for data publishers, but also provides personalized service for data requesters with different utility requirements. We conduct experiments on one simulated dataset and two real datasets, and the simulation results show that DRAPE efficiently achieves a guaranteed privacy level while simultaneously effectively preserving data utility.
Multi-institution credit data sharing-aggregation can improve the accuracy of credit evaluation, however, it also encounters the problems of data fraud, and the privacy, compliance of collaborative modeling. Faced with increasingly stringent and comprehensive privacy protection regulations, the potential security risk call into question the current data sharing-aggregation mode. Blockchain has shown that trusted and auditable computing is possible using a decentralized network of peers, and the immutable distributed ledger. Moreover, oblivious transfer (OT) guarantees the confidentiality of cross-institution transmission and calculation. The integration of blockchain with OT is rapidly increasing in computing environment. In this paper, we design and implement a blockchain-OT-based credit evaluation system dubbed as MEChain. Blockchain acts as a distributed ledger which facilitates efficient and credible data sharing-aggregation. Combined the 1-out-of-N OT, data transformation is used for raw data encryption achieving private, secure, and compliant sharing-aggregation operations. The calculation case and security analysis, prototype system implementation, performance and accuracy rate analysis are presented to validate the proposed solution.
Credit risk refers to the possibility of borrower default, and its assessment is crucial for maintaining financial stability. However, the journey of credit risk data generation is often gradual, and machine learning techniques may not be readily applicable for crafting evaluations at the initial stage of the data accumulation process. This article proposes a credit risk modeling methodology, TED-NN, that first constructs an indicator system based on expert experience, assigns initial weights to the indicator system using the Analytic Hierarchy Process, and then constructs a neural network model based on the indicator system to achieve a smooth transition from an empirical model to a data-driven model. TED-NN can automatically adapt to the gradual accumulation of data, which effectively solves the problem of risk modeling and the smooth transition from no to sufficient data. The effectiveness of this methodology is validated through a specific case of credit risk assessment. Experimental results on a real-world dataset demonstrate that, in the absence of data, the performance of TED-NN is equivalent to the AHP and better than untrained neural networks. As the amount of data increases, TED-NN gradually improves and then surpasses the AHP. When there are sufficient data, its performance approaches that of a fully data-driven neural network model.
Inventory pledge financing (IPF) is a crucial financing way for small and medium-sized enterprise (SMEs). But banks are reluctant to finance SMEs due to fraudulent risk in practice. This paper discusses the application of blockchain in IPF, particularly its impact on mitigating fraud risks. Utilizing game theory models, we illustrate how the finance and operation decisions of participants, along with supply chain efficiency, are influenced by the introduction of blockchain. Meanwhile, equilibrium outcomes are analysed and numerical study is given. Our analysis reveals that under certain conditions, blockchain integration can lead to reduced loan interest rates, lower wholesale prices, increased order quantities by buyers, and enhanced supply chain efficiency. Lastly, we develop a protocol to demonstrate the transfer of digital warehouse receipt on a permissioned blockchain to avoid fraudulent risk. This study provides a theoretical foundation, and a guidance for decisions-making in blockchain-enabled IPF schme.
The Internet of Things (IoT) and distributed ledger technology (DLT) have significantly changed our daily lives. Due to their distributed operational environment and naturally decentralized applications, the convergence of these two technologies indicates a more lavish arrangement for the future. This article develops a comprehensive survey to investigate and illustrate state-of-the-art DLT for various IoT use cases, from smart homes to autonomous vehicles and smart cities. We develop a novel framework for conducting a systematic and comprehensive review of DLT over the IoT by extending the knowledge graph approach. With relevant insights from this review, we extract innovative and pragmatic techniques to DLT design that enable high-performance, sustainable, and highly scalable IoT systems. Our findings support designing an end-to-end IoT-native DLT architecture for the future that fully coordinates network-assisted functionalities.
A key step in cooperative decision-making is for all participants to achieve a consensus that avoids individual favoritism. To reach a consensus, a quantitative systematic mechanism is sometimes preferred. An example of such a mechanism is ranking aggregation, where the task is to rank elements in a certain order. While participating in ranking activities, it is also critical under certain circumstances to protect each decision maker's preference. A promising privacy-preserving technique that is suitable for such a need is differential privacy (DP), which ensures plausible deniability of the protected information with rigorous mathematical guarantee and adjustable privacy level. A concern of the standard DP model is its assumption of letting a curator collect and analyze sensitive information, where in practical situations such a trusted independent curator may not exist. This article proposed a mechanism to solve the above issue using the distributed DP (DDP) framework. The proposed mechanism collects locally differential private rankings from individuals, and then randomly permutes pairwise rankings using a shuffle model to further amplify privacy protection. The final representative is produced by hierarchical rank aggregation (HRA). The mechanism was theoretically analyzed and experimentally compared against the existing methods and demonstrated competitive results in both output accuracy and privacy protection.
Drifting in the direction of earnings surprises for a prolonged period is a decades-puzzling financial anomaly, i.e., the “post-earnings-announcement drift” (PEAD). This paper provided a new simple but effective measure of earnings surprise called ORJ. Based on ORJ, not only is the medium effect of investors' attention on the relationship between earnings surprises and PEAD analyzed but also a tractable and profitable investing strategy is provided. Through comprehensive empirical analysis of the Chinese stock market, we found that i) both earnings surprises and investor attention can increase the degree of PEAD; ii) “good” (bad) earnings surprises strengthen (weaken) the degree of drift by accumulating (decreasing) investor attention; but it is asymmetric that the positive effects of “good” earnings surprises are stronger than the negative effects of “bad” earnings surprises on PEAD; and iii) the strategy obtains an average 6.78% return per quarter in excess of the market but only needs longing dozens of stocks; iv) Typical pricing factors such as the Fama-French three factors, illiquidity and company characteristics have little explanatory power to the returns of the strategy. This paper strongly demonstrates the importance of monitoring overnight returns of earnings announcements to digging the unexpected information and the potential profitability of PEAD in the Chinese market.
Credit risk is the most significant risk faced by credit businesses. Currently, various approaches are widely used in credit evaluation. However, methods based on expert knowledge exhibit obvious subjective cognitive bias, while both statistical and machine learning methods require a substantial amount of historical data. In cases with limited data, the machine-learning effect is poor. Inspired by the structural similarity between neural networks (NN) and the Analytic Hierarchy Process (AHP), we propose a knowledge-augmented dynamic neural network model called KADNN to construct an effective credit evaluation model. This composite architecture will help effectively utilize existing data to alleviate the initial low-data dilemma and can be further utilized for training neural networks. Subsequent data updates can be dynamically incorporated to improve model accuracy. Additionally, this approach improves the comprehensibility and premature convergence issues of the NN model. The proposed approach is validated and evaluated through credit evaluation simulation.
Secure, compliant and authentic multiparty data sharing and collaborative modelling are of great significance to the accuracy of credit evaluation systems. Homomorphic encryption has the feature of supporting ciphertext calculation without sacrificing the accuracy of the model. However, the credibility and security of the existing centralized data homomorphic encryption sharing-aggregation mode have brought great hidden dangers and exacerbated the risk of private data being disclosed. Ensuring the authenticity and controllability of data are also difficulties faced by homomorphic encryption technology. To solve these problems, we propose a novel decentralized privacy-preserving credit evaluation system with trustworthy data content and calculations named PEvaChain based on Hyperledger Fabric blockchain. The PEvaChain consists of three main components: identity management, off-chain encrypted data uploading, and on-chain data security sharing-aggregation. With the help of Hyperledger Fabric's special member access mechanism, incorporating the ciphertext-policy attribute-based encryption (CP-ABE) access control scheme avoids unauthorized access. The original data are transformed by invertible random matrices off the chain, which meets data transfer agreements requirements when data uploading and eliminates the privacy disclosure concerns of data providers to a certain extent. Paillier homomorphic encryption-based data security sharing-aggregation on the chain ensures the security of multiparty data sharing and aggregation while realizing the minimum output and utilization of the original data. Security analysis demonstrates the security and compliance of PEvaChain in terms of data access, encrypted data uploading, sharing-aggregation, and storage. The experimental results show that the proposed approach is feasible, safe and efficient.
Mobile Edge Computing (MEC) has been a promising paradigm for communicating and edge processing of data on the move. We aim to employ Federated Learning (FL) and prominent features of blockchain into MEC architecture such as connected autonomous vehicles to enable complete decentralization, immutability, and rewarding mechanisms simultaneously. FL is advantageous for mobile devices with constrained connectivity since it requires model updates to be delivered to a central point instead of substantial amounts of data communication. For instance, FL in autonomous, connected vehicles can increase data diversity and allow model customization, and predictions are possible even when the vehicles are not connected (by exploiting their local models) for short times. However, existing synchronous FL and Blockchain incur extremely high communication costs due to mobility-induced impairments and do not apply directly to MEC networks. We propose a fully asynchronous Blockchained Federated Learning (BFL) framework referred to as BFL-MEC, in which the mobile clients and their models evolve independently yet guarantee stability in the global learning process. More importantly, we employ post-quantum secure features over BFL-MEC to verify the client's identity and defend against malicious attacks. All of our design assumptions and results are evaluated with extensive simulations.
Increasing the accuracy of prediction models in financial markets is an important but difficult task due to the natural complexities of financial time series, which are nonlinear and nonstationary. This challenge has made machine learning methods popular in recent years. However, the noise contained in financial series dramatically distorts the performance of such approaches. This paper proposes an adaptive denoising method called MIC-EMD that is data-driven and removes the noise contained in input features based on the nonlinear relationship between the target variable and the input features. To verify the advantages of MIC-EMD, a simulation experiment is conducted to compare the performance of several representative denoising methods with that of MIC-EMD. Finally, a comprehensive empirical analysis is performed for the trend predictions of six major indexes in Asian markets using three prevalent machine learning methods (SVM, random forest and LightBGM). After the input features are denoised by MIC-EMD, the results reveal the following: (i) its prediction performance outperforms that of the three learning models with input features denoised by state-of-the-art denoising methods such as ICA, WT, WF, EMD, aEMD, P-EMD and S-EMD, e.g., the prediction accuracies of the three machine learning models increase by 9.77%, 13.5% and 12.3%; and (ii) we obtain a prediction accuracy as high as 70%.
K-lines are among the most popular technical indicators, with considerable attention from investors and frequent use as a reference for investing. This paper provides a novel investing strategy based on a complex network that is built according to the correlations between K-lines. Its monthly return reaches as high as 5.4% for CSI 300 index constituents, significantly outperforming the market. We also found that i) based on a complex network, it takes only 10 to 20 stocks to construct a portfolio that can track the market and ii) portfolios constructed from edge nodes surpass those constructed from high-centrality nodes.
直方图是一种被广为应用的统计数据发布形式,其潜在的隐私泄露风险是当前数据隐私保护领域的关注点.该文针对流数据的直方图发布问题,提出一种符合差分隐私保护要求的方法.其主要特点包括:(1)将w-事件引入流数据的直方图发布加噪机制以确保其满足差分隐私保护需求;(2)采用卡尔曼滤波方式对加噪后的流数据进行后置处理以改善数据效用;(3)通过指数平滑法改进卡尔曼滤波方式避免相邻数据之间的突变性.论文以UCI的两个真实数据集为基础进行流数据直方图模拟发布实验,结果表明该文方法在不同差分隐私预算约束、不同窗口大小情形下均具有明显优势,可在相同隐私保护水平下获得更高的数据可用性.
Financial service quality is private information to the factor (the creator) with her own quality standard, but unknown or unfamiliar to some suppliers (backers) who need to sell their accounts receivable. Credit token transferred from the core enterprise to the upstream multi-tier suppliers can help all suppliers (backers) get equally competitive interest rate due to the endorsement by core enterprise's credit. To maintain a sustainable and healthy platform economy, the most significant issue is incentive for the high-quality financial service factor's (creator's) ongoing engagement. This paper investigates the signaling strategy among a factor (creator) and some suppliers (backers) who need to sell accounts receivable by a game theoretic model in signaling theory. The research finds that the factor (the creator) shall set a higher factoring amount target to signal her high-quality financial service, even if the target is distorted comparing with the optimal level under full information. A separating equilibrium always exists while a pooling equilibrium only exists in some specific conditions. It also shows that high factoring amount comes with high-quality financial service to suppliers (backers) on the platform. This research provides a solution for factors (creators) to offer financial service in the blockchain platform and design campaign in an effectively strategic way. The analysis also displays how information asymmetry affect factors (creators) and suppliers' (backers) decisions, which is closely related with interest rate and has a guidance for platform designers to invite high-quality service factors (creators) to the platform.
Data sharing among different institutions can enhance the integrity of information to evaluate the credit level of financing enterprises more accurately, but it also faces the problems of privacy leakage and data fraud. To solve these problems, this paper proposes a complete linear credit evaluation system with secure sharing and multiparty collaborative computing based on Hyperledger Fabric Blockchain named LEChain. The system consists of four modules: distributed ledger, access control, security aggregation and model storage. With the help of 1-out-of-N oblivious transmission protocol, incorporating elementary matrix transformation data security aggregation realizes the privacy of data transmission and collaborative computation. Finally, a security analysis is given to show the privacy and compliance of the data aggregation.