Clustering populations of networks while recovering their latent hierarchical organization is a fundamental yet largely unexplored problem in network analysis. To formalize this, we introduce the Hierarchical Distance Matrix, a specific class of population-level distance matrices that encodes latent hierarchical organization through recursively nested distance separation, accommodating unbalanced tree depths. Building on this framework, we propose a fully data-driven top-down procedure: network hierarchical clustering based on two-sample testing (NHC-TST). The algorithm recursively splits networks via spectral clustering and uses a graph-based two-sample stopping rule. The procedure adaptively determines the branching structure without requiring prior knowledge of the number of clusters or tree depth. Theoretically, we establish exact recovery of the population-level hierarchical structure and statistical consistency in the empirical procedure. Simulation studies demonstrate highly accurate recovery of both cluster memberships and hierarchical relationships across a wide range of settings. Applied to a global migration dataset, NHC-TST uncovers interpretable multi-resolution temporal structures that are not revealed by conventional flat clustering approaches.
Neutrophils, an essential innate immune cell type with a short lifespan, rely on continuous replenishment from bone marrow (BM) precursors. Although it is established that neutrophils are derived from the granulocyte-macrophage progenitor (GMP), the molecular regulators involved in the differentiation process remain poorly understood. Here we developed a random forest-based machine-learning pipeline, NeuRGI (Neutrophil Regulatory Gene Identifier), which utilized Positive-Unlabeled Learning (PU-learning) and neural network-based in silico gene knockout to identify neutrophil regulators. We interrogated features including gene expression dynamics, physiological characteristics, pathological relatedness, and gene conservation for the model training. Our identified pipeline leads to identifying Mitogen-Activated Protein Kinase-4 (MAP4K4) as a novel neutrophil differentiation regulator. The loss of MAP4K4 in hematopoietic stem cells and progenitors in mice induced neutropenia and impeded the differentiation of neutrophils in the bone marrow. By modulating the phosphorylation level of proteins involved in cell apoptosis, such as STAT5A, MAP4K4 delicately regulates cell apoptosis during the process of neutrophil differentiation. Our work presents a novel regulatory mechanism in neutrophil differentiation and provides a robust prediction model that can be applied to other cellular differentiation processes.
Large Language Models (LLMs) can enhance the performance of Named Entity Recognition (NER) tasks by leveraging external knowledge through in-context learning. When it comes to entity-type-related external knowledge, existing methods mainly provide LLMs with semantic information such as the definition and annotation guidelines of an entity type, leaving the effect of orthographic or morphological information on LLM-based NER unexplored. Besides, it is non-trivial to obtain literal patterns written in natural language to serve LLMs. In this work, we propose LiP-NER, an LLM-based NER framework that utilizes Li teral P atterns, the entity-type-related knowledge that directly describes the orthographic and morphological features of entities. We also propose an LLM-based method to automatically acquire literal patterns, which requires only several sample entities rather than any annotation example, thus further reducing human labor. Our extensive experiments suggest that literal patterns can enhance the performance of LLMs in NER tasks. In further analysis, we found that entity types with relatively standardized naming conventions but limited world knowledge in LLMs, as well as entity types with broad and ambiguous names or definitions yet low internal variation among entities, benefit most from our approach. We found that the most effective written literal patterns are (1) detailed in classification, (2) focused on majority cases rather than minorities, and (3) explicit about obvious literal features.
Periodontitis (PD) and obstructive sleep apnea (OSA) are widespread conditions with profound health consequences. Increasing evidence suggests shared pathophysiological mechanisms between PD and OSA, prompting this study to explore their genetic connections using advanced transcriptomic approaches. Gene expression data was obtained from GEO, integrating bulk and single-cell RNA sequencing (scRNA-seq). Differentially expressed genes (DEGs) were identified, and common DEGs were analyzed via protein–protein interaction (PPI) networks and functional enrichment. Machine learning algorithms, including LASSO, SVM-RFE, and Boruta, were used to screen out hub genes. Expression patterns, diagnostic accuracy, and immune infiltration were assessed. Then, the single-cell analysis was utilized to evaluate cell-specific expression and effects of virtual hub gene knockouts. Drug candidates were predicted using the DSigDB database. In total, 37 common DEGs were identified, in which PECAM1, FCER1G, and THY1 were designated as hub genes. The hub genes were significantly upregulated in disease states, achieving high diagnostic accuracy (AUC > 0.85). Immune infiltration profiles showed differences between the disease and control groups, with hub gene expression positively correlated to plasma cells and M0 macrophages abundance. Single-cell annotation mapped hub gene expression to distinct cell types. Virtual hub gene knockouts highlighted disrupted pathways including oxygen transport and DNA double-strand break repair. Candidate drugs, including pergolide and aspirin, were proposed. This study investigates genetic links between PD and OSA, identifying PECAM1, FCER1G, and THY1 as important diagnostic and therapeutic targets. Integrating multi-omics and machine learning provides a comprehensive approach to unravelling disease interplay and advancing treatment strategies.
Host-based intrusion detection systems (HIDS) have been widely acknowledged as an effective approach for detecting and mitigating malicious activities. Among various data sources utilized in HIDS, system call traces have gained significant popularity due to their inherent advantage of providing fine-grained information. Nevertheless, conventional feature extraction techniques relying on system calls tend to overlook the issue of high-dimensional sparse feature space. In this paper, we conduct a theoretical analysis to investigate the underlying causes of the sparsity problem. Subsequently, we propose an anti-sparse theory (anti-ST) as a solution to address this issue. Then, we design a multi-granularity feature extraction method (MGFE), which also meets the prerequisite mathematical conditions of the anti-ST. By applying this method, we effectively reduce the size of the feature space and minimize the number of generated features, thus mitigating sparsity. Furthermore, leveraging this approach, we propose a robust and anti-sparsity host intrusion detection framework, known as the MGFE-based Host Intrusion Detection Framework (BR-HIDF). A series of experiments were conducted to evaluate the proposed framework and compare it with the state-of-the-art method. The results demonstrate that our framework achieves impressive accuracy (97.26%), precision (97.62%), recall (96.85%), and F1 score (97.23%) in the intrusion detection task, surpassing existing frameworks. Moreover, the proposed framework significantly reduces the time overhead by 38.80%, exhibiting the highest AUC value of 0.992. Furthermore, we enhance the robustness of the detection system by integrating host-based and network-based detection, which provides greater flexibility in identifying various types of attacks.
The recommender system recommends items to the users based on their preferences of implicit feedback. However, implicit feedback often contains noise that deviates from the user’s true preferences, thereby influencing the accuracy of the recommendations. The most effective denoising method is Self-Guided Denoising Learning(SGDL), providing a general denoising scheme that can be applied to various recommendation models. However, it is typically hard to capture user preferences efficiently and enhance the diversity of recommendations to mitigate the ‘filter bubble’ phenomenon. To address these challenges, we propose a novel Data Collaborative Contrastive Recommendation model with self-adaptive noise (DCCR). Specifically, we design the informative item extraction module to mine informative items from the original interactions and improve the accuracy and diversity of recommendations by collaborative training of the informative and the original dataset to learn diverse user embeddings and adapt to noise. Extensive experiments on three public datasets demonstrate the superiority of our DCCR over state-of-the-art methods, balancing the diversity of recommendation lists while optimizing recommendation accuracy.
Autism spectrum disorder (ASD) affects 1-2% of all children and poses a great social and economic challenge for the globe. As a highly heterogeneous neurodevelopmental disorder, the development of its treatment is extremely challenging. Multiple pathways have been linked to the pathogenesis of ASD, including signaling involved in synaptic function, oxytocinergic activities, immune homeostasis, chromatin modifications, and mitochondrial functions. Here, we identify secretagogin (SCGN), a regulator of synaptic transmission, as a new risk gene for ASD. Two heterozygous loss-of-function mutations in SCGN are presented in ASD probands. Deletion of Scgn in zebrafish or mice leads to autism-like behaviors and impairs brain development. Mechanistically, Scgn deficiency disrupts the oxytocin signaling and abnormally activates inflammation in both animal models. Both ASD probands carrying Scgn mutations also show reduced oxytocin levels. Importantly, we demonstrate that the administration of oxytocin and anti-inflammatory drugs can attenuate ASD-associated defects caused by SCGN deficiency. Altogether, we identify a convergence between a potential autism genetic risk factor SCGN, and the pathological deregulation in oxytocinergic signaling and immune responses, providing potential treatment for ASD patients suffering from SCGN deficiency. Our study also indicates that it is critical to identify and stratify ASD patient populations based on their disease mechanisms, which could greatly enhance therapeutic success.
Slot filling is a crucial sub-task in the field of Spoken Language Understanding and aims to match the corresponding semantic slot for each word in the sequence. Slot prediction in an unknown domain requires a large amount of data in the domain for training, but in reality, there is often a lack of trainable samples in the unknown domain, which makes it difficult for the model to predict new domains. This is the biggest challenge of the cross-domain slot filling task. In recent years, the idea of transfer learning has been applied to cross-domain slot filling tasks. The current training method directly mixes the source domain data samples without considering the differences between the various domains in the source domain, which ignores the domain-invariant features contained in the source domain. In this paper, we proposed a cross-domain slot filling model based on multi-domain adaptation. First, we used the domain-adaptive domain projection layer to let the feature learner classify the domain-invariant information and domain-exclusive information into the specified dimension part of the vector, so as to realize the extraction of domain-invariant feature information, and then used the trainable linear transformation matrix to relieve the generalization burden of the feature learner. Experimental results show that our proposed models significantly outperform other methods on average F1-score.
Background and Objective: Most studies used neural activities evoked by linguistic stimuli such as phrases or sentences to decode the language structure. However, compared to linguistic stimuli, it is more common for the human brain to perceive the outside world through non-linguistic stimuli such as natural images, so only relying on linguistic stimuli cannot fully understand the information perceived by the human brain. To address this, an end-to-end mapping model between visual neural activities evoked by non-linguistic stimuli and visual contents is demanded. Methods: Inspired by the success of the Transformer network in neural machine translation and the convolutional neural network (CNN) in computer vision, here a CNN-Transformer hybrid language decoding model is constructed in an end-to-end fashion to decode functional magnetic resonance imaging (fMRI) signals evoked by natural images into descriptive texts about the visual stimuli. Specifically, this model first encodes a semantic sequence extracted by a two-layer 1D CNN from the multi-time visual neural activity into a multi-level abstract representation, then decodes this representation, step by step, into an English sentence. Results: Experimental results show that the decoded texts are semantically consistent with the corresponding ground truth annotations. Additionally, by varying the encoding and decoding layers and modifying the original positional encoding of the Transformer, we found that a specific architecture of the Transformer is required in this work. Conclusions: The study results indicate that the proposed model can decode the visual neural activities evoked by natural images into descriptive text about the visual stimuli in the form of sentences. Hence, it may be considered as a potential computer-aided tool for neuroscientists to understand the neural mechanism of visual information processing in the human brain in the future. (C) 2021 Elsevier B.V. All rights reserved.
Cross-site scripting (XSS) attack is one of the most serious security problems in web applications. Although deep neural network (DNN) has been used in XSS attack detection and achieved unprecedented success, it is vulnerable to adversarial example attacks because its input-output mapping is quite discontinuous to a large extent. The existence of adversarial examples have raised concerns in applying deep learning to key security fields. Therefore, to evaluate the effectiveness of these detection methods, a XSS adversarial example attack technique using Soft Actor-Critic (SAC) reinforcement learning algorithm is presented in the paper. A key aspect of our idea is to train an agent using SAC algorithm to build adversarial examples for several popular XSS detection models which have been proved can achieve very high accuracy rate by simulation experiments. We first design mutation strategies for different modules of XSS attack vectors to ensure the validity of the generated adversarial examples. Then, the agent selects an appropriate escape strategy according to the feedback of the detection model until it bypasses the detection model. The final experiment results show that our model can achieve an escape rate of more than 92% and outperforms the latest method by up to 6%. In other words, the effectiveness of these detection models needs to be improved, at least in terms of defense adversarial example attacks.
The actions of somatostatin (SST) in the nervous system are mediated by specific high affinity SST receptors (SSTR1-5). However, the role of this hormone and the distribution of its receptor subtypes have not yet been defined in neural structures of the human fetus. We have analyzed four neural tissues (CNS, hypothalamus, pituitary and spinal cord) from early to midgestation for the expression of five human SSTR mRNAs, using a reverse transcription–polymerase chain reaction and Southern blot approach. These fetal neural tissues all express mRNA for multiple SSTR subtypes from as early as 16 weeks of fetal life but the developmental patterns of expression vary considerably. Transcripts for SSTR1 and SSTR2A are the most widely distributed, being expressed in all four neural tissues. SSTR2A is often the earliest transcript to be detected (7.5 weeks in CNS). SSTR3 mRNA is confined to the pituitary, hypothalamus, and spinal cord. SSTR4 is expressed in fetal brain, hypothalamus and spinal cord but not pituitary. SSTR5 mRNA is detectable in the pituitary and spinal cord by 14–16 weeks of fetal life. This mapping of SSTR mRNA expression patterns in human fetal neural tissues is an important first step toward our goal of determining the role of SST in the nervous system during early stages in human development.
Aspect-based sentiment analysis aims to identify the aspects mentioned in sentences and their sentiment polarity, which is an important task in fine-grained sentiment analysis. The existing studies use sequence labeling or span-based classification methods, having their own defects such as polarity inconsistency resulted from separately tagging tokens in the former and the heterogeneous categorization in the latter where aspect-related and polarity-related labels are mixed. At the same time, the existing methods ignore the correlation between aspect-polarity pairs in sentences. In order to remedy the above defects, inspiring from the recent advancements in relation extraction, we propose to generate aspect-polarity pairs directly from a text with relation extraction technology, regarding aspect-pairs as unary relations where aspects are entities and the corresponding polarities are relations and utilize sequence decoding to capture the correlation between aspect-polar pairs. The experiments performed on three benchmark datasets demonstrate that our model outperforms the existing state-of-the-art approaches.
Prediction of antimicrobial resistance based on whole-genome sequencing data has attracted greater attention due to its rapidity and convenience. Numerous machine learning-based studies have used genetic variants to predict drug resistance in Mycobacterium tuberculosis (MTB), assuming that variants are homogeneous, and most of these studies, however, have ignored the essential correlation between variants and corresponding genes when encoding variants, and used a limited number of variants as prediction input. In this study, taking advantage of genome-wide variants for drug-resistance prediction and inspired by natural language processing, we summarize drug resistance prediction into document classification, in which variants are considered as words, mutated genes in an isolate as sentences, and an isolate as a document. We propose a novel hierarchical attentive neural network model (HANN) that helps discover drug resistance-related genes and variants and acquire more interpretable biological results. It captures the interaction among variants in a mutated gene as well as among mutated genes in an isolate. Our results show that for the four first-line drugs of isoniazid (INH), rifampicin (RIF), ethambutol (EMB) and pyrazinamide (PZA), the HANN achieves the optimal area under the ROC curve of 97.90, 99.05, 96.44 and 95.14% and the optimal sensitivity of 94.63, 96.31, 92.56 and 87.05%, respectively. In addition, without any domain knowledge, the model identifies drug resistance-related genes and variants consistent with those confirmed by previous studies, and more importantly, it discovers one more potential drug-resistance-related gene.
Ancient literature of Traditional Chinese Medicine (TCM) contains rich clinical experiences, which is the empirical summary of clinical diagnosis and treatment in the process of ancient Chinese medicine practice, and embodies the theoretical framework and ideological basis of the formation and development of TCM. However, due to the volume and dispersion of valuable clinical experiences, it is difficult for TCM doctors to quickly and comprehensively obtain the clinical information they need from ancient literature manually, and the document retrieval tools can only provide documentlevel information screening, which cannot support finegrained information extraction. In addition, the different characteristics of ancient Chinese relative to modern Chinese also limit the use of mainstream text analysis tools. For this reason, we propose a task of information extraction from the ancient literature of TCM for obtaining clinical experiences, which is used to identify text fragments describing clinical experiences in ancient literature and manually annotate sample data for training and testing the extraction task, a sequence labeling model is designed based on deep learning to complete the task. Considering the overfitting problem that can be brought about by the small amount of annotated data, we introduce adversarial training and virtual adversarial training to enhance the generalization ability of the proposed model. A series of sufficient experiments are conducted on the clinical experience dataset to verify the effectiveness of the model, and the experimental results show the feasibility of extracting clinical experiences from ancient literature by information extraction technology, and a promising baseline and a reusable annotated dataset for the new information extraction task are available.
Abstract Abstractive summarization obtains text semantic embedding based on the content of the source corpus to generate summaries, helping users quickly extract valid information from massive amounts of text data. The existing model suffers from poor ability to perceive the correctness of sentence formulation, resulting in the generation of summaries with distorted factual content and representations that do not match the information in the original text, i.e. factual inconsistencies. To address this issue, in this paper, the textual implication task is learned jointly with the summary task, starting from the original text and guided by template sentences. The knowledge of implicit reasoning is incorporated into the encoder model of the summary through parameter sharing to improve the model's ability to perceive facts. First, to enhance the factual consistency of the generated content, this paper uses the extractive summary algorithm Lead-3 to extract partial sentences from the original text as template sentences. Second, we construct a weakly supervised textual implication discriminator dataset on the CNN/Daily Mail. By sharing the same BERT-based text encoder, the text-implication task is jointly trained with the text-summarization task to make the encoder text-implication aware. Finally, experiments were conducted on CNN/Daily Mail to verify the accuracy and factual consistency of the generated summaries by using ROUGE metrics and BERTScore metrics. The results demonstrate that the fact-aware abstractive summarization model proposed in this paper can generate high-quality and factually consistent summaries.
Emoticon, as an emerging network graphic language, is widely used on the social platform due to its ability to express the sentiment and attitude of users intuitively. The current studies take emoticons as text features so that they can neither capture more finegrained correlations between emoticons, nor can they adapt to the development and change of emoticons. In order to overcome the above difficulties, we propose an emoticonimagefeature learning method based on Convolutional Auto Encoder (CAE) for microblog sentiment classification. Our model can learn image features of emoticons by CAE automatically, and such features are incorporated into the embedding representations of microblogs for sentiment classification. We verify the effectiveness of our proposed model on Chinese microblog and twitter datasets, respectively. The experimental results demonstrate that our model outperforms the stateofart methods, and the image features learned by our proposed model have stronger generalization ability even with new emoticons in crosslanguage environment.
When we view a scene, the visual cortex extracts and processes visual information in the scene through various kinds of neural activities. Previous studies have decoded the neural activity into single/multiple semantic category tags which can caption the scene to some extent. However, these tags are isolated words with no grammatical structure, insufficiently conveying what the scene contains. It is well-known that textual language (sentences/phrases) is superior to single word in disclosing the meaning of images as well as reflecting people's real understanding of the images. Here, based on artificial intelligence technologies, we attempted to build a dual-channel language decoding model (DC-LDM) to decode the neural activities evoked by images into language (phrases or short sentences). The DC-LDM consisted of five modules, namely, Image-Extractor, Image-Encoder, Nerve-Extractor, Nerve-Encoder, and Language-Decoder. In addition, we employed a strategy of progressive transfer to train the DC-LDM for improving the performance of language decoding. The results showed that the texts decoded by DC-LDM could describe natural image stimuli accurately and vividly. We adopted six indexes to quantitatively evaluate the difference between the decoded texts and the annotated texts of corresponding visual images, and found that Word2vec-Cosine similarity (WCS) was the best indicator to reflect the similarity between the decoded and the annotated texts. In addition, among different visual cortices, we found that the text decoded by the higher visual cortex was more consistent with the description of the natural image than the lower one. Our decoding model may provide enlightenment in language-based brain-computer interface explorations.