The formal semantics of blockchain smart contracts are the foundation of formal verification. They can be used to establish formal models to verify the security of contracts and help developers understand the specific execution rules of contracts. However, the mathematical logic involved in such modeling poses a high barrier to entry and cannot be directly integrated with other program analysis methods. This article proposes a semantic graph generation approach, KSG, for blockchain smart contracts. First, the semantic rules of the contract language are formally defined, and a semantic interpreter and prover are constructed to automatically transform smart contract code into a scalable semantic graph. This graph incorporates semantic control flow information, semantic data flow information, execution rules, and verification constraints. Next, the generated semantic graph can be utilized for vulnerability detection and symbolic execution and supports iterative optimization based on the analysis results. Finally, the detailed process of semantic graph generation and analysis is demonstrated through the verification of the reentrancy contract and the honeypot contract.
Facilitated with high-frequency observations, we introduce a remarkably parsimonious one-factor volatility model that offers a novel perspective for comprehending daily volatilities of a large number of stocks. Specifically, we propose a multiplicative volatility factor (MVF) model, where stock daily variance is represented by a common variance factor and a multiplicative idiosyncratic component. We demonstrate compelling empirical evidence supporting our model and provide statistical properties for two simple estimation methods. The MVF model reflects important properties of volatilities, applies to both individual stocks and portfolios, can be easily estimated, and leads to exceptional predictive performance in both US stocks and global equity indices.
We establish a framework to study the factor structure in stock variance under a high-frequency and high-dimensional setup. We prove the consistency of conducting principal component analysis on realized variances in estimating the factor structure. Moreover, based on strong empirical evidence, we propose a multiplicative volatility factor (MVF) model, where stock variance is represented by a common variance factor and a multiplicative lognormal idiosyncratic component. We further show that our MVF model leads to significantly improved volatility prediction. The favorable performance of the proposed MVF model is seen in both US stocks and global equity indices.
We study the estimation of high-dimensional covariance matrices under elliptical factor models with 2 + ϵth moment. For such heavy-tailed data, robust estimators like the Huber-type estimator in Fan, Liu and Wang (2018) can not achieve sub-Gaussian convergence rate. In this paper, we develop an idiosyncratic-projected self-normalization (IPSN) method to remove the effect of heavy-tailed scalar parameter, and propose a robust pilot estimator for the scatter matrix that achieves the sub-Gaussian rate. We further develop an estimator of the covariance matrix and show that it achieves a faster convergence rate than the generic POET estimator in Fan, Liu and Wang (2018).
A composite smart contract can execute smart contracts that may belong to other owners or companies through external calls, bringing more security challenges to blockchain applications. Traditional static verification methods are inadequate for analyzing the dynamic execution of these contracts, resulting in misjudgment and omission issues. Therefore, this paper proposes a model checking approach based on dynamic behavior that verifies the security and business logic of composite smart contracts. Utilizing automata, the method models contracts, users, attackers, and extracts properties, focusing on six types of common security vulnerabilities. A thorough case study and experimental evaluation demonstrate the method’s efficiency in identifying vulnerabilities and ensuring alignment with business requirements. The UPPAAL tool is employed for comprehensive verification, proving its effectiveness in enhancing smart contract security.
We propose a Degree-Corrected Block Model with Dependent Multivariate Poisson edges (DCBM-DMP) to study stock co-jump dependence. To estimate the community structure, we extend the SCORE algorithm in Jin (2015) and develop a Spectral Clustering On Ratios-of-Eigenvectors for networks with Dependent Multivariate Poisson edges (SCORE-DMP) algorithm. We prove that SCORE-DMP enjoys strong consistency in community detection. Empirically, using high-frequency data of S&P 500 constituents, we construct two co-jump networks according to whether the market jumps and find that they exhibit different community features than GICS. We further show that the co-jump networks help in stock return prediction.
We establish a high-dimensional statistical learning framework for individualized asset allocation. Our proposed methodology addresses continuous-action decision-making with a large number of characteristics. We develop a discretization approach to model the effect of continuous actions and allow the discretization frequency to be large and diverge with the number of observations. The value function of continuous-action is estimated using penalized regression with our proposed generalized penalties that are imposed on linear transformations of the model coefficients. We show that our proposed Discretization and Regression with generalized fOlded concaVe penalty on Effect discontinuity (DROVE) approach enjoys desirable theoretical properties and allows for statistical inference of the optimal value associated with optimal decision-making. Empirically, the proposed framework is exercised with the Health and Retirement Study data in finding individualized optimal asset allocation. The results show that our individualized optimal strategy improves the population financial well-being.
We study the estimation of the high-dimensional covariance matrix andits eigenvalues under dynamic volatility models. Data under such modelshave nonlinear dependency both cross-sectionally and temporally. We firstinvestigate the empirical spectral distribution (ESD) of the sample covariancematrix under scalar BEKK models and establish conditions under which thelimiting spectral distribution (LSD) is either the same as or different fromthe i.i.d. case. We then propose a time-variation adjusted (TV-adj) sample co-variance matrix and prove that its LSD follows the same Marcenko-Pasturlaw as the i.i.d. case. Based on the asymptotics of the TV-adj sample co-variance matrix, we develop a consistent population spectrum estimator and an asymptotically optimal nonlinear shrinkage estimator of the unconditionalcovariance matrix
Deep Neural Networks (DNNs) have shown exceptional promise in providing Artificial Intelligence (AI) to many computer vision applications. Nevertheless, complex models, intensive computations, and resource constraints hinder the use of DNNs in edge computing scenarios. Existing studies have already focused on edge-cloud collaborative inference, where a DNN model is decoupled at an intermediate layer, and the two parts are sequentially executed at the edge device and the cloud server, respectively. In this work, we examine the status quo approaches of DNN execution, and find that it still takes a lot of time on edge device computation and edge-cloud data transmission. Using this insight, we propose a new edge-cloud DNN collaborative computing framework, JMDC, based on Joint Model and Data Compression. In JMDC, we adopt the attention mechanism to select important model channels for efficient inference computation, and important regions of the intermediate output for transferring to the cloud. We further use the quantization technique to reduce actual bits needed to be transferred. Depending on the specific application requirements on latency or accuracy, JMDC can adaptively determine the optimal partition point and compression strategy under different resource conditions. By extensive experiments based on the Raspberry Pi 4B device and the CIFAR10 dataset, we demonstrate the effectiveness of JMDC in enabling on-demand low-latency DNN inference, and its superiority over other baseline schemes.
There are various benefits of providing undergraduate research experience at higher education institutions, such as increasing students' interest and retention, improving critical thinking, and building self-confidence. It also has some challenges such as the lack of time and academic maturity. Course-based Undergraduate Research Experience (CURE) not only inherits the benefits of undergraduate research experience, but also effectively avoids the challenges. In this paper, we report on a CURE project in information security, focusing on vulnerability scanning. Surveys were conducted for assessment and the results indicate that the research project could help participants better understand how to do research. The participants felt more confident in conducting research and their problem-solving skills were improved.
在当前移动互联网时代,数据量增长迅速,服务计算能力不断增强,数据隐私保护和服务环境可信成为备受关注的重要问题.本文研究面向卷积神经网络典型应用场景的可信隐私服务计算模型,探索支持同态加密的数据和模型计算方法,保护数据隐私.构建基于区块链和智能合约技术服务过程存证及计算权益分配方法,保证服务计算的公开透明、可信可追溯.探索资源提供者、模型拥有者及用户的新型云环境资源数据服务模式,促进资源有效整合,发展共享经济.最后,通过实验分析该模型的隐私保护方法.
BackgroundSuperconducting undulator (SCU) prototype with small magnet gap of 5 mm, long magnet length of 4 m and high magnet field of 1.58 T was being developed at Shanghai High Repetition rate XFEL and Extreme light facility (SHINE). Compared to any other superconducting undulator, there is no cryocooler being installed on the cryostat in this SCU prototype.PurposeThis study aims at the cooling design for the binary current leads for SCU's normal operating.MethodsBinary current leads composed of normal conductive copper leads and high temperature superconducting current leads (HTS) were adopted for SCU to connect superconducting coils inside the cryostat and outer cables. Low-temperature helium gas was used to transport independent refrigerator system to the cooling tubes inside the prototype, hence the binary current leads were cooled. Thermal conduction components installed on the middle of the thermal shield were employed to transfer heat load of normal conductive copper leads, and heat load of copper leads was optimized by simulation. Auxiliary superconducting rods were designed for connecting cold ends of HTS in the cryostat test.ResultsThe temperature difference between hot ends of HTS and low-temperature helium gas is less than 20 K from the result of cryostat test, all binary current leads is operating normally with full current.ConclusionsIt is practicable to use cooling tubes with low-temperature helium gas to cool binary current leads of the SCU prototype by thermal conduction, which is different from cooling solution for current leads in any other SCU being developed presently.
As machine learning methods become more powerful and capture more nuances of human behavior, biases in the dataset can shape what the model learns and is evaluated on. This paper explores and attempts to quantify the uncertainties and biases due to annotator demographics when creating sentiment analysis datasets. We ask >1000 crowdworkers to provide their demographic information and annotations for multimodal sentiment data and its component modalities. We show that demographic differences among annotators impute a significant effect on their ratings, and that these effects also occur in each component modality. We compare predictions of different state-of-the-art multimodal machine learning algorithms against annotations provided by different demographic groups, and find that changing annotator demographics can cause >4.5 in accuracy difference when determining positive versus negative sentiment. Our findings underscore the importance of accounting for crowdworker attributes, such as demographics, when building datasets, evaluating algorithms, and interpreting results for sentiment analysis.
Deep Neural Networks (DNNs) have achieved remarkable success in many computer vision tasks recently, but the huge number of parameters and the high computation overhead hinder their deployments on resource-constrained edge devices. It is worth noting that channel pruning is an effective approach for compressing DNN models. A critical challenge is to determine which channels are to be removed, so that the model accuracy will not be negatively affected. In this paper, we first propose Spatial and Channel Attention (SCA), a new attention module combining both spatial and channel attention that respectively focuses on "where" and "what" are the most informative parts. Guided by the scale values generated by SCA for measuring channel importance, we further propose a new channel pruning approach called Channel Pruning guided by Spatial and Channel Attention (CPSCA). Experimental results indicate that SCA achieves the best inference accuracy, while incurring negligibly extra resource consumption, compared to other state-of-the-art attention modules. Our evaluation on two benchmark datasets shows that, with the guidance of SCA, our CPSCA approach achieves higher inference accuracy than other state-of-the-art pruning methods under the same pruning ratios.
Modern machine learning algorithms typically require large amounts of labeled training data to fit a reliable model. To minimize the cost of data collection, researchers often employ techniques such as crowdsourcing and web scraping. However, web data and human annotations are known to exhibit high margins of error, resulting in sizable amounts of incorrect labels. Poorly labeled training data can cause models to overfit to the noise distribution, crippling performance in real-world applications. In this work, we investigate the viability of using data augmentation in conjunction with semi-supervised learning to improve the label noise robustness of image classification models. We conduct several experiments using noisy variants of the CIFAR-10 image classification dataset to benchmark our method against existing algorithms. Experimental results show that our augmentative SSL approach improves upon the state-of-the-art.
Unmanned Aerial Vehicle (UAV) can play an important role in wireless systems as it can be deployed flexibly to help improve coverage and quality of communication. In this paper, we consider a UAV-assisted Mobile Edge Computing (MEC) system, in which a UAV equipped with computing resources can provide offloading services to nearby user equipments (UEs). The UE offloads a portion of the computing tasks to the UAV, while the remaining tasks are locally executed at this UE. Subject to constraints on discrete variables and energy consumption, we aim to minimize the maximum processing delay by jointly optimizing user scheduling, task offloading ratio, UAV flight angle and flight speed. Considering the non-convexity of this problem, the high-dimensional state space and the continuous action space, we propose a computation offloading algorithm based on Deep Deterministic Policy Gradient (DDPG) in Reinforcement Learning (RL). With this algorithm, we can obtain the optimal computation offloading policy in an uncontrollable dynamic environment. Extensive experiments have been conducted, and the results show that the proposed DDPG-based algorithm can quickly converge to the optimum. Meanwhile, our algorithm can achieve a significant improvement in processing delay as compared with baseline algorithms, e.g., Deep Q Network (DQN).
Multimodal classification is a core task in human-centric machine learning. We observe that information is highly complementary across modalities, thus unimodal information can be drastically sparsified prior to multimodal fusion without loss of accuracy. To this end, we present Sparse Fusion Transformers (SFT), a novel multimodal fusion method for transformers that performs comparably to existing state-of-the-art methods while having greatly reduced memory footprint and computation cost. Key to our idea is a sparse-pooling block that reduces unimodal token sets prior to cross-modality modeling. Evaluations are conducted on multiple multimodal benchmark datasets for a wide range of classification tasks. State-of-the-art performance is obtained on multiple benchmarks under similar experiment conditions, while reporting up to six-fold reduction in computational cost and memory requirements. Extensive ablation studies showcase our benefits of combining sparsification and multimodal learning over naive approaches. This paves the way for enabling multimodal learning on low-resource devices.
Data and Analysis of Data (or Data Analytics, or DA) are becoming increasingly important parts of professional practices within IT (Information Technology) and across the larger domains of business and science. Many dedicated data analytics related programs have been developed to train the future data analyst workforce. Still, growing needs of DA skills and knowledge almost as literacy requirement for many job practices outpace what the dedicated DA programs can supply. DA training often goes beyond dedicated DA curricula. Even in IT field, many non-DA IT courses could involve some DA topics. Currently, there is a lack of study on how to coordinate the pedagogical methods in teaching those DA topics embedded in different courses to deliver needed DA training for our non-DA IT graduates. Here, we explore the possibility of developing a pedagogy framework that helps link the teaching and learning of those fragmented DA topics embedded across different non-DA courses in multidisciplinary study areas of IT field together. In doing so, we expect this interdisciplinary effort to provide a complimentary approach to promote DA education and prepare IT graduates with the DA skills needed in a variety of job requirements. Although our current effort is limited to the IT field this experience can inspire and be transferred to other fields where DA is a part of the prevalent job practices for their DA pedagogy development.