
Despite the broad application of deep reinforcement learning (RL), transferring and adapting the policy to unseen but similar environments is still a significant challenge. Recently, the language-conditioned policy is proposed to facilitate policy transfer through learning the joint representation of observation and text that catches the compact and invariant information across environments. Existing studies of language-conditioned RL methods often learn the joint representation as a simple latent layer for the given instances (episode-specific observation and text), which inevitably includes noisy or irrelevant information and cause spurious correlations that are dependent on instances, thus hurting generalization performance and training efficiency. To address this issue, we propose a conceptual reinforcement learning (CRL) framework to learn the concept-like joint representation for language-conditioned policy. The key insight is that concepts are compact and invariant representations in human cognition through extracting similarities from numerous instances in real-world. In CRL, we propose a multi-level attention encoder and two mutual information constraints for learning compact and invariant concepts. Verified in two challenging environments, RTFM and Messenger, CRL significantly improves the training efficiency (up to 70%) and generalization ability (up to 30%) to the new environment dynamics.
Geo-distributed ML training can benefit many emerging ML scenarios (e.g., large model training, federated learning) with multi-regional cloud resources and wide area network. However, its efficiency is limited due to 2 challenges. First, efficient elastic scheduling of multi-regional cloud resources is usually missing, affecting resource utilization and performance of training. Second, training communication on WAN is still the main overhead, easily subjected to low bandwidth and high fluctuations of WAN. In this paper, we propose a framework, Cloudless-Training, to realize efficient PS-based geo-distributed ML training in 3 aspects. First, it uses a two-layer architecture with control and physical training planes to support elastic scheduling and communication for multi-regional clouds in a serverless maner.Second, it provides an elastic scheduling strategy that can deploy training workflows adaptively according to the heterogeneity of available cloud resources and distribution of pre-existing training datasets. Third, it provides 2 new synchronization strategies for training partitions among clouds, including asynchronous SGD with gradient accumulation (ASGD-GA) and inter-PS model averaging (MA). It is implemented with OpenFaaS and evaluated on Tencent Cloud. Experiments show that Cloudless-Training can support general ML training in a geo-distributed way, greatly improve resource utilization (e.g., 9.2%-24.0% training cost reduction) and synchronization efficiency (e.g., 1.7x training speedup over baseline at most) with model correctness guarantees.
In order to study the effect of solid particles of different shapes in the gas pipelines on the erosion wear charac-teristics of the elbow and predict the wear distribution of the inner wall of the elbow after being impacted,computa-tional fluid dynamics(CFD)is used to calculate the gas flow.The discrete element method(DEM)is used to cal-culate the particle motion,and the non-spherical particles are described by the super-ellipsoids model.The shear impact energy model(SIEM)is used to calculate elbow wear.The simulation results show that,under the same gas velocity,although the difference of the movement of particles of different shapes in the elbow is not evident,the corresponding wear condition of the inner wall of the elbow is significantly different from each other:the closer the particle shape is to the square,the higher the maximum wear rate in the elbow,and the wider the range of serious wear.But the position of the maximum wear rate will tend to be stable.The probability of sliding friction of prolate ellipsoid particles in the elbow is higher,while the probability of rolling friction of oblate ellipsoid particles is high-er,which causes the total wear rate of prolate ellipsoid particles on the elbow to be higher than that of the oblate el-lipsoid.Comparing the influence of particle shape on the wear of the elbow under different gas velocities,it can be found that larger gas velocity only homogeneously increases the wear of the elbow and the maximum erosion zone will not change.
In this paper,the shock vibration characteristic and the fatigue life of the sampling mechanism of sludge drying moisture detection device are investigated based on the discrete element method and finite element method(DEM-FEM).The impact load distribution under different particle sizes and the impact characteristic under different re-ceiving angles are analyzed.Furthermore,the modal vibration patterns and amplitude-frequency responses under different excitations are studied,and the fatigue life distribution and safety factor distribution under different loads are analyzed.The results show that particle size and receiving angle have significant influence on impact character-istics,and the avoidance between the motor frequency and the structural inherent frequency during the advance of the sampling mechanism reaches 24%.Under normal light loads,the number of stress cycles that can be withstood is 1×108,while the minimum safety factor is approximately 0.54 under a heavy load of 100 N.The fatigue failure and fracture tend to occur at the connecting rods.
In this paper,numerical artificial boundary conditions for the damped dispersive wave problems are designed.The matching boundary condition method is used to construct accurate local artificial boundary conditions by matc-hing the characteristic frequency-wave number relationship of damped dispersive waves.The matching boundary conditions take a linear combination form of atomic displacement and velocity near the artificial boundary,and the combination coefficients are determined by matching the frequency-wave number relationship.In this work,a one-dimensional infinitely long damped monoatomic chain is taken as an example,and the frequency-wave number rela-tionship of damped dispersive waves is established with frequency rather than wave number as the independent vari-able.The proposed matching boundary conditions can effectively deal with damped waves with different wave speeds and different spatial attenuation rates.Both reflection coefficient analysis and numerical examples verify the validity of the artificial boundary conditions.Matching boundary conditions are compact in form,low cost in compu-tation,and efficient in absorption,and can be applied to molecular dynamics simulation and multi-scale calculation of crystals.
It is hard to implement multi-table Hash join on hardware accelerators.On one hand,multi-table join has an indefinite number of tables and various connection modes.The flexibility in multi-table Hash join is in contradiction with fixed hardware architectures.On the other hand,the capacity of intermediate results expands with the number of tables increasing.The capability of data management and monitoring asks for higher hardware overhead.To ena-ble flexible and efficient multi-table Hash join,a software-hardware co-optimization methodology is proposed.Soft-ware subsystem abstracts multi-table Hash join into forward and reverse computation modes,and agilely organizes Hash join processes.Additionally,the memory access and computing are collaboratively optimized in hardware de-sign.A regular hardware Hash table is designed to improve memory bandwidth.Meanwhile,a homogeneous com-puting engine is designed to perform both forward and reverse computation.To further improve the efficiency of Hash join,multi data channels and an instruction control system are configured.The experiment results showed that a single computing engine could improve the performance of multi-table Hash join 9.2-11.0 times higher than con-tral processing unit(CPU).Furthermore,the 8-way parallel multi-table Hash join engines could make full use of DDR bandwidth resources and get 71.1 times performance of CPU.
How to determine the layout of static/const data is a big challenge faced by tensor program automatic genera-tion frameworks.Ansor,the most broadly-used and promising framework among them,solves this issue by training a performance cost model according to a layout strategy specified in advance,then searching the tensor program with the optimal performance based on the cost model.However,there are two problems:a single strategy cannot be suitable for all tasks,and the performance cost model is not accurate.In order to solve these problems,AL-An-sor,a tensor program automatic generation framework based on the adaptive layout(AL)strategy of static data,is proposed.It adaptively chooses multiple layout strategies during the search process,and trains the performance cost model according to them.In this way,AL-Ansor can find a tensor program with higher performance.Taking convo-lutional layers as workloads,this work evaluates Ansor and AL-Ansor in a target server with a 32-core Intel Xeon CPU.The experimental results show that AL-Ansor improves the execution performance by 13.81%,12.41%,and 16.59%,respectively,on average,compared against Ansor with three specified layout strategies.
Regulatory rules detection in policy text is an emerging natural language processing task,which has important application value for policy conflict detection,policy intelligent retrieval,regulatory compliance inspection,and e-government system requirements engineering.This paper takes the detection of mineral resources regulatory rules as the research goal,and proposes a detection method based on bidirectional encoder representation from transform-ers(BERT)with prompts.By constructing a prompt template with[MASK],which incorporates regulatory rules information,the proposed method can give full play to the auto-encoding advantages of the mask language model,effectively stimulate BERT model to extract text features related to regulatory rules and increase the stability of the model.A new application mode of regulatory rules detection based on BERT model is proposed,which uses the[MASK]hidden vector instead of the[CLS]hidden vector for classification and prediction.The experimental re-sults on the dataset of mineral resources regulatory rules show that the accuracy,macro-average F1 score and weigh-ted-average F1 score of this method are better than the baseline methods.The experimental results on the public dataset also show the effectiveness of the proposed method.
The high randomness and complexity of arc fault make it difficult to be accurately identified.Aiming at the problem that the traditional arc recognition algorithm has low real-time performance and high hardware computing power,an error minimization extreme learning machine(EM-ELM)arc fault detection method suitable for edge computing,multi-load types and multi-feature combination is proposed.Through fast Fourier transform(FFT)and db4 wavelet decomposition,the period mean difference,pulse width percentage,inter-harmonic factor and wavelet high-frequency energy are extracted as the input characteristics of the arc fault detection algorithm on the edge side.On this basis,OS-EM-ELM combined with online sequence(OS)method is proposed,and the algo-rithm is improved by using field operation data to improve adaptability.The experimental results show that the pro-posed edge side arc fault detection method can effectively distinguish the normal and arc fault waveform,and it is suitable for the complex situation of working with a variety of loads at the same time.The calculation amount is small,the real-time performance is high,the adaptability is strong,and the application cost is low,which is more in line with the requirements of edge calculation of arc detection device.
The existing radio signal modulation identification methods are usually difficult to effectively identify the un-classified signal when the prior data is insufficient.To solve this problem,this paper proposes a deep transfer clus-tering(DTC)of radio signals method based on knowledge transfer.This method analyzes the similarity between samples based on sample comparison,and uses a convolutional neural network(CNN)to extract the features of ra-dio signals.At the same time,a pre-training framework is designed,which effectively improves the feature extrac-tion ability of CNN by transferring the knowledge of the same domain dataset and achieves the goal of guiding the clustering direction and improving the clustering performance.The experimental results show that the clustering per-formance of this method is significantly better than the existing clustering methods on multiple public datasets.Compared with existing methods,the clustering accuracy of DTC on the RML2016.10A and RML2016.04C data-sets is improved by 30.34%and 28.04%,respectively.
Aiming at the digital implementation of resonant clock network in the integrated circuit design,this paper pro-poses a modeling and optimization method of resonant clock circuits(MRC),which simplifies the integration process of resonant clock networks.At present,traditional simulation tools for building resonant circuit models is time consuming,and the existing resonant circuit models cannot meet the requirements of rapid implementation and digital library construction.According to the three-stage circuit state of the resonant designs,the polyline reduction model in this paper can obtain the current waveforms of various resonant circuits quickly and accurately.An optimi-zation objective function of global power consumption is also given based on this model,providing a theoretical basis for the selection of circuit parameters.The post-Spice simulation results based on 12 nm Fin-FET technology show that the model accuracy is more than 90%and can accurately fit the actual power consumption trend.Matlab-based implementation of the proposed model can achieve 105 times speedup compared with Spice-based simulation.
In order to improve the coordination of human-robot cooperation,this paper proposes an adaptive system of hu-man-robot interaction based on non-zero-sum game.The system consists of inner loop and outer loop which are de-coupled.In the outer loop,the human-robot cooperation control is designed by introducing a non-zero-sum game method,and an energy function about human and robot force is constructed,and the optimal control is achieved by solving the Nash equilibrium.The neural network estimator is employed to update the uncertain parameters in the energy function and estimate the output force of human and robot.By designing the central value of the neural net-work function,the relationship between the robot force and the tracking error is obtained to ensure the tracking per-formance of the method.During the update process,the stiffness coefficient is adaptively adjusted when external force exists,so as to realize the compliance and coordination of human and robot.In addition,a neural network controller is designed in the inner loop,and the radial basis neural network is applied to approximate the unknown robot dynamics model using the collected input and output data of robot system,which improves the tracking accu-racy of the system.Simulation results verify the effectiveness of this method.
3D-HEVC标准中引入了具有大面积平坦区域、陡峭边缘和低纹理复杂度特性的深度图.针对深度图编码过程中编码单元(CU)率失真优化导致编码复杂度过高这一问题,本文在分析深度图编码所具有的特点的基础上,构建了深度图划分深度数据集,并提出了一种基于两通道特征传递卷积神经网络(T-CNN)的划分深度预测算法.使用本文提出的算法替换原始编码器中各视点下深度图CU划分模块,可以在一定的率失真性能损失下,将原始HTM-16.0 编码器编码时间平均减少76%左右,编码效率得到了显著提升.
伴随5G标准的不断演进和商用网络的规模部署,5G已成为引领我国智能制造高质量发展的新引擎.与此同时,以高带宽、高频次小包通信为特征的工业应用也对 5G终端基带芯片协议处理提出了挑战.本文提出一种以数据面加速器(DPA)为核心的高性能软硬件协同5G协议处理架构,该架构将异构芯片计算资源与协议处理功能进行了合理映射,并通过并行化设计大幅提升5G用户面数据处理性能.实验结果表明,相比纯软件的实现方案本文提出的协同架构在不同业务负载条件下,数据包处理时延平均下降28.3%,包处理通量平均提升38%.在0.5ms的时隙周期配置下,本文架构的数据包处理速率大于2000 包/s,可以满足工业5G大规模现场节点集中式数据采集的需求.
区块链是分布式的数据存储系统,共识算法为区块链实现安全存储数据提供支撑和保障.Kafka作为共识算法中的一种,其高吞吐速率、低时延的特点受到青睐.但使用Kafka算法的系统接受大量交易时,易产生数据倾斜,即分布式系统的多节点结构中,大量数据集中在少数节点,导致系统资源被占用、性能下降.为解决上述问题,本文提出基于时间序列模型长短期记忆网络(LSTM)的智能优化方法.通过学习过往生产者接收到的交易量,预测下一时刻面临的交易量,动态调整生产者节点数量,减少数据集中在少数节点的情况.实验结果显示,本文方法可以将Kafka系统时延降低 2~3 倍,吞吐速率提升2~3 倍,与优化前相比系统效率提升52.62%,比2 种传统优化方法分别提升近3%和40%,能耗仅小幅提升,系统使用情况保持更加合理.
电空制动是轨道车辆应用最广泛的黏着制动方式,其制动性能主要受制于轮轨间的黏着状态.在复杂低黏着条件下,传统制动控制系统面临的最大问题是无法使黏着时刻保持最优.因此,基于轮轨黏滑特性和车辆动力学理论,本文首先建立以黏着观测器为核心的蠕滑寻优模型;其次提出以Levenberg Marquardt(L-M)算法为核心的神经网络控制器,完成最优黏着控制系统;最后使用Matlab/Simulink平台分别对基于多交替轨面和实验低黏着轨面的列车黏着控制进行仿真模拟,并与传统比例积分微分控制器(PID)作用下的黏着情况做对比.结果表明,即使面对具有不同特性的低黏着轨面,轮轨黏着在控制系统作用下都能迅速维持在当前轨面下的最优值,有效缩短了制动距离和时间.相比传统PID控制,本文提出的控制系统在制动时间和制动距离上同比减小 4.9%与4.1%,调控能力更强,适用于低黏着和大蠕滑下的列车制动工况.
针对当前远端内存系统中页面预取与页面替换因操作系统与应用程序之间语义鸿沟导致的局限性问题,本文提出一个软硬件协同的远端内存系统.通过在内存控制器中增加热点页面提取表,将实时访存的热点页面信息通过内存中的缓冲区传送给操作系统.同时,通过对访存信息的学习,构建了高精度的异步预取框架与替换框架,降低应用关键数据路径的开销,提升远端内存系统的性能.本文利用内存跟踪工具构建了一个原型仿真系统.实验证明,在拥有全局实时访存信息后,预取框架可以实现超过 90%的准确率与覆盖率,与谷歌提出的远端内存系统Fastswap相比,性能提升 59%.相比于内核默认替换框架,替换框架使应用性能提升30%.
含裂纹正交各向异性圆柱壳自由振动响应及其裂纹形态求解是一个复杂的结构动力学问题.针对此问题,提出一种基于线弹簧模型和波传播方法的含斜裂纹正交各向异性圆柱壳的裂纹形态求解方法.基于Kirchhoff-Love壳体理论,建立含斜裂纹的正交各向异性圆柱壳力学模型,得到经典边界条件下的正交各向异性圆柱壳的自由振动响应特性.利用线弹簧模型计算裂纹区域的局部柔度,构造了不同裂纹形态的附加应力关系.结合波传播方法,获得含斜裂纹的正交各向异性圆柱壳的自由振动响应特性,进而得到一种基于固有频率的裂纹形态识别方法.研究结果表明,裂纹的存在会导致圆柱壳局部柔度的降低和固有频率的下降,且裂纹的尺寸越大、深度越深,固有频率的下降程度越大;随着裂纹角度的增大,固有频率的下降程度先增大后减小;通过不同裂纹形态产生的固有频率变化规律,可以对裂纹形态进行识别,得到裂纹的几何空间分布.研究结果可为正交各向异性薄壳裂纹形态求解和振动响应求解方面的研究提供有益参考,也可为圆柱壳结构的裂纹损伤识别方面提供理论支持.
卷积神经网络传统的应用平台是中央处理器(CPU)和图形处理器(GPU),其体积和功耗不能适应轻量化的行业,轻量化的专用集成电路(ASIC)平台专用加速器的开发成本又不能适应愈发复杂和深层次的网络结构.针对上述问题,设计一种基于现场可编程门阵列(FPGA)的卷积神经网络(CNN)加速器,既满足轻量化应用场景,又有低开发成本的特性.设计浮点加法器和浮点乘法器组合成卷积运算的基本运算单元,完成 16 bits浮点数乘累加操作只需要消耗一个数字信号处理器(DSP)资源;针对FPGA运算特性设计了基于ReLU函数的激活层模块;设计可调节并行度的各层模块,可根据平台资源在性能、功耗和面积上取得平衡;设计用比较器简化的SoftMax模块.实验结果表明,在 100 MHz工作频率下,峰值算力可达44.8 GFLOPS,功率仅为4.51 W.
针对协同过滤推荐算法中用户所交互的物品对其决策的不同贡献度问题,提出了一种基于相关注意力的协同过滤推荐算法.该算法结合深度学习中的注意力机制为不同物品分配不同的权值来捕获与目标物品最相关的物品,探索不同物品的权重对模型预测的影响并以此提升推荐的准确度;在此基础上,为了解决推荐算法鲁棒性低的问题,进一步提出了注意力协同对抗性训练的推荐算法,通过对抗性学习的方法并使用快速梯度符号算法(FGSM)构建对抗样本输入模型进行对抗训练,缓解模型受扰动的影响从而提升算法鲁棒性.在Pinterest和MovieLens-1M这2 个数据集上的实验结果表明,所提算法不仅有效提升了推荐算法的准确度,同时也增强了推荐系统的鲁棒性.