Pathogenic bioaerosols are critical for outbreaks of airborne disease; however, rapidly and accurately identifying pathogens directly from complex air environments remains highly challenging. We present an advanced method that combines open-set deep learning (OSDL) with single-cell Raman spectroscopy to identify pathogens in real-world air containing diverse unknown indigenous bacteria that cannot be fully included in training sets. To test and further enhance identification, we constructed the Raman datasets of aerosolized bacteria. Through optimizing OSDL algorithms and training strategies, Raman-OSDL achieves 93% accuracy for five target airborne pathogens, 84% accuracy for untrained air bacteria, and 36% reduction in false positive rates compared to conventional close-set algorithms. It offers a high detection sensitivity down to 1:1000. When applied to real air containing >4600 bacterial species, our method accurately identifies single or multiple pathogens simultaneously within an hour. This single-cell tool advances rapidly surveilling pathogens in complex environments to prevent infection transmission.
In recent years, with the widening applications of the Internet of Things (IoT), more and more perception services (e.g. air quality indicator services, road traffic congestion monitoring services, etc) with different arguments (e.g. data type, source location, creator, etc) will be deployed by dedicated IT infrastructure service providers for constructing customized IoT systems with low cost by subscription. So it is an indispensable step to check whether the required perception services with specified arguments have been available for the constructing IoT through discovery method to reduce the redundancy of service deployment. However, it is a challenging problem to design efficient (i.e. achieving high accuracy and low response delay with low overhead), highly robust, and trustworthy mechanisms for discovering perception services on resource-constrained IoT devices. To solve this problem, we proposed a distributed service discovery method, named VSA-SD, based on the Vector Symbolic Architecture (VSA). This method employs hyperdimensional vectors to describe services in a distributed manner, and measures the degree of service matching by calculating the Hamming distance, thereby achieving service discovery. We implemented VSA-SD in NBUFlow, which is an IoT task construction and offloading test platform, and evaluated its performance through comprehensive experiments. Results show that VSA-SD outperforms the centralized, hybrid, and other distributed service discovery mechanisms in terms of accuracy, response delay, overhead, robustness, trustability, interoperability, and mobility.
心律不齐是一种常见的心脏疾病,严重时可能会危及生命,因此对该疾病开展早期筛查和分类在临床医学中具有重要意义.搭载心电信号(ECG)传感器的可穿戴设备凭借低成本和便捷等特点,是实现日常心脏健康监测的理想平台之一.然而受制于计算能力等因素的限制,可穿戴设备需要将数据上传到云端进行分析,增加了等待时延和用户隐私泄露风险.另一方面,现有心律不齐分类算法在训练时受疾病样本分布不平衡等因素的影响,在识别部分异常病症时的表现不尽人意,限制了其应用范围.为解决上述问题,本文提出了一种基于压缩卷积神经网络的心律不齐分类算法,增强了其在移动平台上的部署能力.同时在训练过程中通过将类别先验分布引入损失函数中,提升了算法对异常病症的识别能力.实验结果表明,本文提出的压缩模型相比经典模型在减少98.2%参数量的同时,超越了许多相关工作取得了0.759 的宏F1 值.
Electrocardiogram(ECG) reconstruction is studied to rebuild ECG from heterogeneous biosignal, such as Photoplethys-mography(PPG), to combine the diagnosis experience of ECG and the convenient collection of PPG. Heterogeneous biosignals are information-lost channels compared with ECG, so simple mapping from the heterogeneous biosignals to ECG leads to unsatisfactory performance. In this paper, we present an invertible neural network, called IR-ECG(invertible reconstruction of ECG), to model the processes of ECG reconstruction. We deliberately design the invertible block, called TimeFlow, to construct the flow-based model. In the forward process, IR-ECG produces the heterogeneous biosignal and captures the distribution of lost information from ECG. In the backward process, ECG reconstruction is finished using a randomly-drawn latent vector and heterogeneous biosignal. Experimental results show that IR-ECG outperforms existing works in quantitative, visual and semantic evaluations.
Dynamic Anomaly Detection (DAD) in building energy consumption is crucial for enhancing energy efficiency. Two main categories of abnormal energy consumption sources, equipment faults and operation and maintenance (O&M) issues, are involved. While technology like the Internet of Things (IoT) can detect equipment faults, research on abnormal energy consumption related to O&M issues is limited. In this study, a DAD method was developed, which combines a white-box model for energy consumption related to O&M issues with IoT for real-time data input. The optimal inspection interval was determined using a confusion matrix and precision-recall curve. To quantify preferences for false detections and missed detections, a customized confidence probability and real-time confidence interval were introduced based on a precedence chart and questionnaire survey. Using air-conditioning energy consumption in a Shanghai office building as an example, the optimal inspection interval for four O&M cases was found to be 1 hour. Similar tolerance levels for missed and false detections were observed among building O&M personnel. The optimal confidence probability was found to be 77.2%, resulting in a DAD precision of 92.5% and a recall of 84.1%, confirming the validity of the method. The impact of O&M preferences on anomaly detection metrics was also examined, and the 77.2% confidence probability was found to be suitable for various practical engineering applications. These results provide a universal prior conclusion for the proposed preference methodology for weighing missed and false detections in anomaly detection.
现有基于深度学习的水下声呐图像目标检测方法受限于水下声呐图像噪声大、信噪比低,因而检测精度有限.针对该问题,本文提出了基于投影感知和声呐参数信息嵌入的水下声呐图像目标检测方法SonarNet.提出的非参数化的投影感知对齐模块(PAA)在不引入额外的训练参数且无需额外标注的情况下,通过提取水下目标的投影区域特征与目标本身特征融合来提升目标检测精度.同时为了提升算法在不同声呐工作参数下的鲁棒性,本文设计了一个轻量级的声呐全连接网络SonarMLP,将声呐设备的工作参数信息以嵌入信息的形式引入到目标检测过程中.本文在声呐图像目标检测数据集上对算法的有效性进行了验证,在有效检测出水下目标的同时,比现有常用深度学习方法有更高的检测精度,能够提升3%以上的各类平均精确度(mAP).
Breast cancer is one of the common malignant tumors in women. It seriously endangers women's life and health. The human epidermal growth factor receptor 2 (HER2) protein is responsible for the division and growth of healthy breast cells. The overexpression of the HER2 protein is generally evaluated by immunohistochemistry (IHC). The IHC evaluation criteria mainly includes three indexes: staining intensity, circumferential membrane staining pattern, and proportion of positive cells. Manually scoring HER2 IHC images is an error-prone, variable, and time-consuming work. To solve these problems, this study proposes an automated predictive method for scoring whole-slide images (WSI) of HER2 slides based on a deep learning network. A total of 95 HER2 pathological slides from September 2021 to December 2021 were included. The average patch level precision and f1 score were 95.77% and 83.09%, respectively. The overall accuracy of automated scoring for slide-level classification was 97.9%. The proposed method showed excellent specificity for all IHC 0 and 3+ slides and most 1+ and 2+ slides. The evaluation effect of the integrated method is better than the effect of using the staining result only.
The automatic diagnosis of arrhythmia using machine learning has been a hot topic and extensively researched recently. A common problem is class imbalance that could make the deep learning models easily trapped into biased learning towards the majority class while ignoring rare classes during reasoning. When conducting inter-patient experiments, the inherent individual difference makes the condition even worse. Current deep learning methods generally take elaborate data modification strategies like data augmentation that complicate the training process. This paper, however, presents a special Hybrid Convolutional Transformer Network (HCTNet) that could effectively extract decisive patterns by drawing on doctors’ diagnosis experience in structure design. Meanwhile, a novel logit adjusted loss is applied to enlarge the pairwise margin between different classes so that the HCTNet could be highly sensitive to rare anomalies. In the experiments, the proposed method has outperformed most state-of-the-arts on the benchmark of the MIT-BIH database: the F1 scores for the three primary arrhythmias (N, S, V) are 97.5%, 61.5%, and 88.3%, respectively under the inter-patient paradigm.
Electrocardiogram(ECG) is commonly utilized in clinical diagnosis and health monitoring. However, ECG acquisition is cumbersome because ECG electrodes must be attached to the skin tightly. Ballistocardiogram(BCG), originating from the heartbeat, can be captured without attaching the BCG sensors to the skin, making BCG collecting convenient and non-feeling. However, BCG diagnostic experience is limited compared with ECG. In this paper, we propose a model, called Rec-AUNet, to reconstruct ECG from BCG so that we can combine the convenient measurement of BCG and diagnostic experience on ECG. Rec-AUNet utilizes an attentive UNet-based neural network with encoding paths to better capture the temporal features and with attentive paths to preserve the spatial features. To better evaluate the coherence of the global waveform and the fidelity of the local physiological features between the synthetic ECG and the original ECG, we deliberately design the person identification task as the semantic metric. The synthetic ECG from our proposed model achieves an accuracy of 93.75% and 94.05% in person identification on the Kansas-Dataset and selfcollected data, respectively, outperforming existing algorithms.
在诸多物联网实际应用中,原始采集信号数据多含有大量噪声,特别是在运动相关场景里.需从含大量噪声的一维时序信号中对有效信号活动区域起止点进行准确识别,以支持相关分析.已有的基于双阈值规则的识别方法对噪声十分敏感,噪声的存在会导致计算出的识别阈值无法匹配非噪声段的原始数据,从而导致将随机噪声数据识别为信号活动区间或者漏检信号活动区间.基于机器学习和深度学习的识别方法需要大量的样本数据,在样本量较小的物联网场景中模型会产生欠拟合问题,从而降低识别精度.为了对含有大量噪声且数据量少的一维时序信号中的信号活动区间进行准确识别,提出了一种基于局部动态阈值的信号活动区间识别方法EasiLTOM(signal activity interval recognition based on local dynamic threshold).该方法基于局域信号计算识别阈值,并使用最短信号长度对噪声尖峰进行过滤,可避免随机噪声对信号活动区间识别的影响,解决漏检和误检问题,从而提高识别精度.此外,EasiLTOM方法所需数据量小,适用于数据稀少的物联网场景.为验证EasiLTOM方法的有效性,该研究于3个月间采集了 14人次的表面肌电数据,并使用2个公开数据集进行了对比实验.结果表明:EasiLTOM方法对信号活动区间可达到平均93.17%的识别精度,相对于现有的双阈值和机器学习方法,分别提升了 15.03%和4.70%,在运动分析相关场景中具有实用价值.
Deep learning-based methods have recently shown great promise in the defect detection task. However, current methods rely on large-scale annotated data and are unable to adapt a trained deep learning model to new samples that were not observed during training. To address this issue, we propose a new siamese defect-aware attention network (SDANet) with a template comparison detection strategy that improves the defect detection technique for matching new samples without rapidly collecting new data and retraining the model. In SDANet, the siamese feature pyramid network is used to extract multi-scale features from input and template images, the defect-aware attention module is proposed to obtain inconsistency between input and template features and use it to enhance abnormality in input image features, and the self-calibration module is developed to calibrate the alignment error between the input and template features. SDANet can be used as a plug-in module to enable most existing mainstream detection algorithms to detect defects using not only the features of defects, but also the inconsistency between features of the inspected image and the template image. Extensive experiments on two publicly available industrial defect detection benchmarks highlight the effectiveness of our method. SDANet can be seamlessly integrated into mainstream detection methods and improve the mAP of mainstream detection algorithms on unseen samples by 12% on average which outperforms current state-of-the-art method by 7.7%. It can also improve the performance in seen samples by 4.3% on average. SDANet can be used in general defect detection applications of industrial manufacturing.
Objective To propose an intelligent quantitative analysis method of Ki-67 index for breast cancer immunohistochemical whole slide image (WSI). Methods The pathological sections of patients with breast cancer diagnosed and treated in Peking Union Medical College Hospital from January 2020 to December 2020 were retrospectively collected, and scanned at 40 magnification as WSI images. Manual interpretation of the Ki-67 index was conducted by 2 pathologists according to the guidelines formulated by the International Breast Cancer Ki-67 Working Group in 2019, which is considered the gold standard. According to the ratio of 5:8, WSI was randomly divided into two data sets, A and B (data set A was randomly divided into training set, validation set and test set according to a ratio of 7:1:2). After the hot spot area in WSI of the data set A was manually marked, each WSI randomly cropped 2000 512×512 pixel patches in the 40 field of view, and 50 patches of them were randomly selected to label tumor cells and calculate the Ki-67 index. The conditional random field model was used to fuse the spatial features of the image blocks, the features were extracted by the ResNet34 pre-training model to construct a hot spot recognition model, and its performance (accuracy) was evaluated in the test set. In the hot spot area, 10 fields of view were randomly selected under the high-power field of view (×40), and the model automatically completed the cell classification and calculated the average Ki-67 index. Taking the results of manual interpretation as the gold standard, the accuracy of the Ki-67 index evaluation results of the data set B by the model was calculated, and the Bland-Altman method was used to evaluate the consistency between the results of manual interpretation and model analysis. Results A total of 132 pathological sections of patients with breast cancer which met the inclusion and exclusion criteria were selected. There were 50 images in data set A (35, 5, and 10 images in training set, validation set, and test set, including 70 000, 10 000, and 20 000 patches, respectively), and 82 images in data set B. The average accuracy of the model for identifying hot spots in the test set was 81.5%, and the accuracy of the Ki-67 index calculation results for the B data set was 90.2%. Bland-Altman analysis showed that the Ki-67 index calculated by manual interpretation and model was in good agreement. Conclusion The intelligent quantitative analysis method of Ki-67 index proposed in this study has high accuracy and can assist pathologists to achieve efficient interpretation of Ki-67 index.
Transformer models have demonstrated their promising potential and achieved excellent performance on a series of computer vision tasks. However, the huge computational cost of vision transformers hinders their deployment and application to edge devices. Recent works have proposed to find and remove the unimportant units of vision transformers. Despite achieving remarkable results, these methods take one dimension of network width into consideration and ignore network depth, which is another important dimension for pruning vision transformers. Therefore, we propose a Width & Depth Pruning (WDPruning) framework that reduces both width and depth dimensions simultaneously. Specifically, for width pruning, a set of learnable pruning-related parameters is used to adaptively adjust the width of transformer. For depth pruning, we introduce several shallow classifiers by using the intermediate information of the transformer blocks, which allows images to be classified by shallow classifiers instead of the deeper classifiers. In the inference period, all of the blocks after shallow classifiers can be dropped so they don’t bring additional parameters and computation. Experimental results on benchmark datasets demonstrate that the proposed method can significantly reduce the computational costs of mainstream vision transformers such as DeiT and Swin Transformer with a minor accuracy drop. In particular, on ILSVRC-12, we achieve over 22% pruning ratio of FLOPs by compressing DeiT-Base, even with an increase of 0.14% Top-1 accuracy.
Sequential recommendation aims to suggest items to users based on sequential dependencies. Graph neural networks (GNNs) are recently proposed to capture transitions of items by treating session sequences as graph-structured data. However, existing graph construction approaches mainly focus on the directional dependency of items and ignore benefits of feature aggregation from undirectional relationship. In this paper, we innovatively propose a joint graph contextualized network (JGCN) for sequential recommendation, which constructs both the directed graphs and undirected graphs to jointly capture current interests and global preferences. Specifically, we introduce gate graph neural networks and model the combined embedding of weighted position and node information from directed graphs for capturing current interests. Besides, to learn global preferences, we propose a graph collaborative attention network with correlation-based similarity of items from undirected graphs. Finally, a feed-forward layer with the residual connection is applied to synthetically obtain accurate transitions of items. Extensive experiments conducted on three datasets show that JGCN outperforms state-of-the-art methods.
On account of a large scale of dataset need to be annotated to train the deep learning based modern object detection model, zero-shot object detection has become an important research field which aims to simultaneously localize and recognize unseen objects that are not observed during training. In order to improve the performance of zero-shot object detection, recent state of the art methods tend to make complicated modifications to the modern object detectors in terms of the model structure, loss function and training process. They always take the simple modification as a baseline, and think it is worse than more complicated methods. In contrast, we find that simple modification can achieve better performance. Considering that the redundant modification may increase the risk of over-fitting in seen classes and reduce generalization performance on unseen classes, we propose a visual language based succinct zero-shot object detection framework, which only replaces the classification branch in the modern object detector with a lightweight visuallanguage network. Since zero-shot object detection is a classic multi-modal learning protocol which consists of a visual feature space and a language space, our visual-language network learns the visual language alignment from the image and language data of seen classes and transfers this alignment to detect unseen objects. Following the Occam's razor principle that "Entities should not be multiplied unnecessarily", extensive experimental results show that our succinct framework can suppress all existing zero-shot object detection methods on several benchmarks and gets the new state-of-the-art.
Personalized tag recommender systems recommend a series of tags for items by leveraging users' historical records, which helps tag-aware recommender systems (TRS) to better depict user profiles and item characteristics. However, existing personalized tag recommendation solutions are insufficient to capture the collaborative signal hidden in the interactions among entities without considering reasonable correlations, since neighborhood messages are treated as the same weights when constructing graph-structured data, resulting in decreased accuracy in making recommendations. In this paper, we propose a Tag-aware Attentional Graph Neural Network (TA-GNN), which integrates the attention mechanism into tag-based graph neural networks to alleviate the above issues. Specifically, we extract the user-tag interaction and the item-tag interaction from the user-tag-item graph structure. For each interaction, we exploit the contextual semantics of multi-hop neighbors by leveraging attentional strategy on graph neural networks to discriminate the importance of different connected nodes. In this way, we effectively extract collaborative signals of neighborhood representations and capture the potential information in an explicit manner. Extensive experiments on three public datasets show that our proposed TA-GNN outperforms the state-of-the-art personalized tag recommendation baselines.
Factorization Machines (FMs) are extensively used for sparse contextual prediction tasks by modeling feature interactions. Despite successful application of FM and its abundant deep learning variants, these improved FMs mainly focus on capturing feature interaction at the vector-wise level while ignoring the more sophisticated bit-wise information. In this paper, we propose a novel Convolutional Feature-interacted Factorization Machine (CFFM), which learns crucial interactive patterns from enhanced feature-interacted maps with non-linearity. Specifically, in the high-order feature interactions part of CFFM, we propose a special Convolutional Max Pooling (Conv-MP) block to adequately learn interaction patterns from both vector-wise and bit-wise perspectives. Besides, we improve linear regression in FMs by incorporating a linear attention mechanism. Extensive experiments on two public datasets demonstrate that CFFM outperforms several state-of-the-art approaches.
针对物联网控制规则在实际应用中,由于延迟触发执行导致物联网系统引发不同程度的控制安全事故,本文基于规则的实时特性提出了一种面向物联网控制规则的端云动态分配方法(EasiDEP),以提高物联网系统控制规则的实时性.首先,根据规则之间的相关性以及实时特性,提出了一种规则聚类算法,将规则库中的规则分为实时和无实时规则;然后,基于规则聚类结果提出一种端云动态规则分配方法,将实时规则分发至云端和本地规则服务端同时处理,提高实时规则的触发执行率.最后,通过实验验证了Eas-iDEP能够有效地提高物联网控制规则的实时触发执行率,降低系统发生控制安全事故的风险.
目的 探索基于机器学习的人工智能(Artificial Intelligence,AI)辅助诊疗系统在非特异性腰痛诊断中的应用效果.方法 使用Thought Technology Ltd生产的FlexComp Infiniti 10肌电仪采集受试者腰部左右两侧多裂肌、左右两侧最长肌和左右两侧腰髂肋肌的肌电信号,15名受试者均为确诊非特异性腰痛患者,一次采集流程包括屈曲放松运动五次,双足桥式运动、左足桥式运动、右足桥式运动和Biering Sorensen等长运动各一次.3名资深医师使用常规诊断方式对患者进行诊断,并以此为标准,对比评价AI辅助诊断系统获得的结果.结果 AI辅助诊疗系统对本文选取的15名非特异性腰痛患者的检出准确性达到100%,且对于肌肉募集能力、疲劳速度和静息速度的检测结果与医生常规诊断方式相比具有良好的一致性,平均诊断用时减少26.3 min,具有统计学意义(P<0.05).结论 初步验证表明,该AI辅助诊疗系统可对非特异性腰痛提供准确高效的辅助诊断及量化评估,为临床检测提供可靠帮助.