Graph neural networks (GNNs) have achieved remarkable success in unsupervised graph anomaly detection (GAD) tasks. The existing GNN-based unsupervised GAD models typically employ self-supervised learning to capture the intrinsic low-dimensional representations of data, thereby adapting to the inductive bias of GNNs toward homophily. However, they often ignore the anomaly-discriminative property of nodes, termed the one-class homophily property, i.e., normal nodes tend to have strong affinity with each other, while the homophily in anomalous nodes is significantly weaker than that in normal nodes. In this paper, motivated by the one-class homophily, we propose a novel local-affinity-based adaptive graph filter (LAF) to address label imbalance challenge in unsupervised GAD tasks. Our method generates subgraphs by removing heterophilic edges from the raw graph, thereby strengthening the isolation of anomalous nodes and enhancing high-frequency information of the graph structure. Subsequently, we apply an adaptive graph filter to each subgraph to dynamically integrate low-frequency and high-frequency information, so as to better model the anomalous features. Instead of minimizing the commonly used data reconstruction errors, our method optimizes the model by maximizing the local node affinity. The final result is obtained by averaging the anomalous scores from multiple models trained on different subgraphs, during the inference phase. Experimental results on six real-world GAD datasets show that LAF outperforms the existing baseline algorithms in unsupervised GAD tasks.
Live streaming has become a cornerstone of today's internet, enabling massive real-time social interactions. However, it faces severe risks arising from sparse, coordinated malicious behaviors among multiple participants, which are often concealed within normal activities and challenging to detect timely and accurately. In this work, we provide a pioneering study on risk assessment in live streaming rooms, characterized by weak supervision where only room-level labels are available. We formulate the task as a Multiple Instance Learning (MIL) problem, treating each room as a bag and defining structured user-timeslot capsules as instances. These capsules represent subsequences of user actions within specific time windows, encapsulating localized behavioral patterns. Based on this formulation, we propose AC-MIL, an Action-aware Capsule MIL framework that models both individual behaviors and group-level coordination patterns. AC-MIL captures multi-granular semantics and behavioral cues through a serial and parallel architecture that jointly encodes temporal dynamics and cross-user dependencies. These signals are integrated for robust room-level risk prediction, while also offering interpretable evidence at the behavior segment level. Extensive experiments on large-scale industrial datasets from Douyin demonstrate that AC-MIL significantly outperforms MIL and sequential baselines, establishing new state-of-the-art performance in room-level risk assessment for live streaming. Moreover, AC-MIL provides capsule-level interpretability, enabling identification of risky behavior segments as actionable evidence for intervention. The project page is available at: https://qiaoyran.github.io/AC-MIL/.
Anaerobic digestion (AD) of food waste (FW) is a key waste-to-energy strategy, yet daily biogas yield is often challenging to sustain, partly due to a limited understanding of the internal methanogens and their functional divergence. Here, we investigated seven full-scale mesophilic FW-AD systems distributed across China along a broad latitudinal gradient (>2800 km), linking methane production variations (0.38-2.11 m3/m3•d-1) with the phylogenetic distributions of methanogens and their methanogenic genes. We found that hydrogenotrophic and aceticlastic pathways were ubiquitous, whereas methylotrophic methanogenesis showed regional enrichment in warmer regions, reflecting persistent influences of climate-associated upstream conditions on downstream methanogenic communities. Gene-level phylogeny of methanogenesis-related alleles, rather than species-level phylogeny, closely tracked biogas yield variation (Mantel's P < .05) and showed consistently stronger associations than gene-level compositions (mean standardized total effect: 0.491 vs. 0.298, P < .01). Higher methane yields (1.61 vs. 0.61 m3/m3•d-1 in high- vs. low-performing systems, P < .01) were significantly associated with reduced Faith's phylogenetic diversity (1.82 vs. 2.30, P < .01) and tighter clustering (mean pairwise phylogenetic distance: 0.25 vs. 0.30, P < .01) of methanogenic gene variants, suggesting that phylogenetic coherence may reflect ecological filtering favoring efficient methanogenesis, albeit at the expense of functional redundancy. These findings highlight gene-level trait phylogeny as a potential proxy for functional robustness, offering a framework for ecological design of AD microbiomes.
Graph invariant learning aims to acquire invariant node representations across various environments, achieving substantial success in addressing Out-of-Distribution (OOD) generalization for graph learning tasks. As obtaining environment splits on graphs is typically costly, most graph invariant learning methods heavily depend on inferring the underlying environments to learn invariant node representations. Due to the high heterogeneity of graph data without explicit source labels, existing environment inference methods cannot simultaneously satisfy the requirements of diversity and similarity. To address this challenge, we propose an approach called sOft environment inFerence with Test-timE adaptatioN, abbreviated as OFTEN, which enables us to perform graph invariant learning without any predefined environment split or partition information. The intuition is to enhance the diversity among environments while preserving the original graph topology. Extensive experiments on several graph OOD benchmarks demonstrate the consistent superiority of OFTEN across all settings.
Abstract Metagenome sequencing not only plays a pivotal role in unravelling the genetic diversity and functional potential of microbial communities but also facilitates the discovery of genome context for microbial dark matter. This study presents a comparative analysis of metagenome sequencing strategies, focusing on the impact of read length on the assembly quality of metagenome binning. We employed metaSPAdes assembly with varying k‐mer lists and the read lengths on 19 Illumina datasets, revealing that longer reads significantly improve the number of contigs and their length, despite a trade‐off in N50. Specially, longer reads also contribute to better performance of gene fragment reconstruction from contigs. Next, the substantial potential of Nanopore sequencing was further evaluated by comparing the short‐read assembly by Illumina, long‐read assembly by Nanopore and hybrid assembly strategies on samples from extreme environments, including both cold seep and hot spring. The binning of assembled contigs and subsequent metagenome‐assembled genome quality assessment highlighted the superiority of long‐read data in reconstructing medium‐ and high‐quality drafted genomes, specifically, increasing medium‐quality species‐level representative genomes by 1.32‐fold. These findings advocate for the integration of extended read lengths and Nanopore sequencing in metagenome analysis, which can lead to a more nuanced comprehension of the environmental microbiome.
The rise of live streaming has transformed online interaction, enabling massive real-time engagement but also exposing platforms to complex risks such as scams and coordinated malicious behaviors. Detecting these risks is challenging because harmful actions often accumulate gradually and recur across seemingly unrelated streams. To address this, we propose CS-VAR (Cross-Session Evidence-Aware Retrieval-Augmented Detector) for live streaming risk assessment. In CS-VAR, a lightweight, domain-specific model performs fast session-level risk inference, guided during training by a Large Language Model (LLM) that reasons over retrieved cross-session behavioral evidence and transfers its local-to-global insights to the small model. This design enables the small model to recognize recurring patterns across streams, perform structured risk assessment, and maintain efficiency for real-time deployment. Extensive offline experiments on large-scale industrial datasets, combined with online validation, demonstrate the state-of-the-art performance of CS-VAR. Furthermore, CS-VAR provides interpretable, localized signals that effectively empower real-world moderation for live streaming.
Antibiotic resistance genes (ARGs) present in food waste pose a significant environmental and public health challenge, with anaerobic digestion emerging as a promising technology to reduce ARG abundance during waste treatment. In this study, we analyzed the resistomes in 64 anaerobic digestion sludge samples from seven full-scale food waste treatment facilities representing seven Chinese provinces. Across all facilities, a small core set of glycopeptide (van clusters), β-lactamase, aminoglycoside, and macrolide-lincosamide-streptogramin genes accounted for most ARG abundance (70.3%), marking them as critical targets for monitoring and post-treatment at high-risk sites such as Wenzhou. Resistome composition differed significantly among facilities and exhibited moderate correlation with bacterial taxonomic composition, with Firmicutes (Bacillota), Chloroflexota, and Proteobacteria as the major carriers associated with multiple resistance classes. ARG abundance was positively correlated with mobile genetic elements (r = 0.54, p < 0.0001), driven by integrases, transposases, and Tn916. Horizontal gene transfer was largely constrained within phylogenetic boundaries, particularly within Firmicutes (66.67%), limiting cross-phyla ARG dissemination. Resistome variation was driven predominantly by deterministic processes.; these deterministic filters together with regional differences in food-waste composition and MGEs, collectively select for a glycopeptide-dominated, Firmicutes-anchored resistome that is distinct from those in activated sludge and manure digesters.
Chinese Baijiu, a globally esteemed traditional distilled spirit, owes its profound complexity and distinct sensory attributes to a dynamic and complex microbial consortium. While microbial diversity and composition in specific fermentation stages have been studied, a comprehensive, large-scale investigation of how these brewing-suitable microbiomes assemble along the entire environmental gradient of production has been lacking. To bridge this knowledge gap, this study conducted a large-scale investigation across a seven-habitat gradient, spanning peripheral township dust, outdoor and indoor workshop environments, Daqu, and Zaopei, in 12 distinct distilleries in Maotai Town (northwestern Guizhou, China). Our findings revealed marked distinctions in microbial diversity, community composition, and inferred species interactions across the seven habitats, with the fermentation habitats exhibiting lower diversity, enrichment of specific taxa, and simpler yet tighter networks. A significant “dispersal barrier” between indoor and outdoor workshop environments was identified, demonstrating that active, bidirectional microbial immigration occurs primarily within workshops, reciprocally shaping environmental dust and fermentation habitats. Conversely, immigration from outdoor environments was strongly constrained. Importantly, microbial community assembly mechanisms diverged: workshops-indoor habitats (dust and fermentation) were predominantly shaped by homogeneous selection, while workshops-outdoor habitats were primarily governed by heterogeneous selection and dispersal limitation. In summary, the consistently high-temperature, high-humidity conditions of Maotai Town distillery workshops impose strong environmental filtering, driving homogeneous selection to enrich fermentation-adapted microbes while limiting exogenous immigration. This work provided evidence for a semi-permeable ecological framework for high-temperature production settings in Maotai Town and offered actionable insights for distillery management and quality control.
A critical limitation of conventional sequential recommendation models (SRMs) is their reliance on observed user-item interaction sequences within a closed-world setting, which hinders their ability to generalize to unseen or infrequent items. Recently, Large Language Models (LLMs) have shown remarkable promise in recommendation systems due to their vast world knowledge and advanced reasoning capabilities. Current research has predominantly explored two approaches: using LLMs to directly generate recommendations and distilling knowledge from LLMs to enhance conventional SRMs. However, these approaches face two major challenges: (1) high inference costs, as they require LLM responses during inference, either for generating predictions or as supplementary input; (2) inadequate distillation of the reasoning process, as existing methods focus mainly on improving embeddings or aligning outputs, without fully integrating LLMs' inherent reasoning capabilities. To address these issues, we propose LCKD-SR, an LLM-driven Cascaded Knowledge Distillation framework for Sequential Recommendation. In this framework, an LLM, a Teacher SRM, and a Student SRM form a hierarchical distillation structure, enabling an LLM-free inference by using only the Student model. Beyond traditional embedding and ranking distillation, our framework abstracts the LLM's sequential reasoning abilities by identifying key interactions that subsequently guide the Teacher's attention using learnable markers. The Student model, which mirrors the architecture of the Teacher, achieves seamless knowledge alignment from the Teacher across all three aspects. Extensive experiments demonstrate the effectiveness and efficiency of the proposed LCKD-SR, showcasing its scalability to perform multi-level knowledge transfer while enabling LLM-independent inference, thereby overcoming the inference cost and reasoning limitations of existing methods.
Anaerobic digestion (AD) is a cornerstone technology for sustainable waste treatment and renewable energy recovery, yet its complex microbe-metabolite interactions remain poorly understood. Here, we combined high-resolution molecular profiling and microbial community sequencing in a three-month study across seven full-scale digesters to resolve dissolved organic matter (DOM) and microbiome dynamics. A total of 28 925 DOM molecules, including a conserved core of 1154 metabolites, were identified. By disentangling metabolic pathways, we observed complex transformation patterns that extend beyond simple substrate breakdown. Molecules within a mass window (183.57-390.81 m/z) exhibited high persistence, strong microbial associations, and distinct transformation trajectories. Within this mass window, microbial community composition and feedstock input, together explained ~30.1%-43.4% of the observed spatiotemporal variation. In each digester, 1260-2108 molecules were closely associated with microbial metabolism, forming 7.77-24.52 microbe-metabolite associations on average. The accumulation and turnover of these microbial metabolites were strongly linked to methane production and system performance, highlighting microbial processing of DOM as a significant factor shaping microbe-metabolite interactions. This perspective emphasizes the importance of microbe-metabolite interplay in AD, providing a conceptual framework for predictive monitoring and optimization of engineered biotechnologies.
Graph Transformers (GTs), as emerging foundational encoders for graph-structured data, have shown promising performance due to the integration of local graph structures with global attention mechanisms. However, the complex attention functions and their coupling with graph structures incur significant computational overhead, particularly in large-scale graphs. In this paper, we decouple graph structures from Transformers and propose the Graph-Agnostic Linear Transformer (GALiT). In GALiT, graph structures are solely utilized to denoise raw node features before training, as our findings reveal that these denoised features have integrated the main information of the graph structure and can replace it to guide Transformers. By excluding graph structures from the training and inference stages, GALiT serves as a graph-agnostic model which significantly reduces computational complexity. Additionally, we simplify the linear attention functions inherited from traditional Transformers, which further reduces computational overhead while still capturing the relationships between nodes. Through weighted combination, we integrate the denoised features into the attention mechanism, as our theoretical analysis reveals the key role of the synergy between linear attention and denoised features in enhancing representation diversity. Despite decoupling graph structures and simplifying attention mechanisms, our model surprisingly outperforms most GNNs and GTs on benchmark graphs. Experimental results indicate that GALiT achieves high efficiency while maintaining or even enhancing performance.
Understanding the environmental occurrence patterns of soilborne pathogens is essential for public health, yet a comprehensive and accurate assessment remains challenging. This study presents an innovative technical framework integrating metagenomic pathogen screening with quantitative validation using chip-based digital PCR (dPCR) targeting the overall bacteria community as well as three dominant species-Ralstonia pickettii, Saccharomonospora viridis, and Gordonia terrae. This approach enabled a comprehensive quantification of potential human-, plant-, and zoonotic pathogens and elucidation of their environmental drivers across urban soil habitats in Beijing. Farmland and hospital greenspaces exhibited higher potential pathogen richness (15.55 ± 5.87 and 10.70 ± 4.52) and abundance (22,475.52 ± 15,559.92 and 26,217.62 ± 19,299.90 copies g⁻¹ soil) compared with forests and campus greenspaces. The composition of potential pathogens varied among habitats, with farmlands containing the highest number of unique species, and four taxa were detected across all habitats, showing strong adaptive capacity. Pathogen diversity was positively correlated with total and available phosphorus and with total bacterial α- and β-diversity, while negatively associated with soil organic carbon, reflecting limited pathogen inputs in carbon-rich forest soils and the key role of phosphorus in pathogen enrichment. Climatic and soil physicochemical factors indirectly influenced pathogen diversity by modulating bacterial communities, whereas human activities directly increased pathogen abundance. Molecular ecological network analysis demonstrated that 81% of the associations between pathogenic and non-pathogenic taxa were significantly negative, suggesting competitive exclusion as a key regulatory mechanism. Collectively, these findings provide a precise monitoring framework and new insights into cross-species interactions, contributing to improved risk assessment and One Health strategies for the prevention of soilborne diseases.
In this paper, we reveal that existing Differentially Private Graph Neural Networks (DP-GNNs) are not effective against Graph Reconstruction Attack (GRA). We further attribute the ineffectiveness of existing DP-GNNs against GRA to their unstructured perturbation mechanism, which only induces unidirectional shift in the embedding similarity distribution. Specifically, this perturbation mechanism tends to decrease the embedding similarity of all node pairs without significantly disrupting the relative ranking, thus allowing GRA to still reconstruct the original graph structure by leveraging the relative ranking of similarities. To address this, we propose a novel Differentially Private Graph Neural Network based on Structured Perturbation (GRASP). Specifically, we observe that independent noise tends to decrease the embedding similarity, while identical noise tends to increase it. By integrating these two types of noise using a Bernoulli technique, we introduce a simple yet effective structured perturbation mechanism, which promotes bidirectional shift in the embedding similarity distribution, thereby effectively disrupting the relative ranking and defending against GRA. Extensive experiments on eight benchmark datasets demonstrate that GRASP effectively defends against GRA. Furthermore, GRASP achieves a superior privacy-utility trade-off compared to existing graph structure protection methods. The implementation of GRASP is available at https://github.com/ZhiyuZone/GRASP/.
Graph Neural Networks (GNNs) have demonstrated impressive success across diverse fields when data satisfies in-distribution (ID) assumption. Nevertheless, GNN performance significantly declines in cases of distribution shifts between training and testing graph data. This degradation primarily stems from spurious correlations between irrelevant domain information and target labels in out-of-distribution (OOD) scenarios. Thus, maximizing the utilization of domain information becomes imperative. In light of this, we propose a novel approach named Domain-aware Node Representation Learning (DNRL), comprehensively incorporates domain information to bolster generalization capability. Specifically, DNRL selectively interpolates nodes with the same label but different domains, extending training data into unseen domains and alleviating the effects caused by domain-related spurious correlations. Futhermore, by introducing a domain-aware contrastive learning strategy, our method implicitly decouples domain information from node information to learn domain-independent node representations. Extensive experiments on graph out-of-distribution benchmarks demonstrate that DNRL can achieve effective OOD generalization performance across diverse domains.
Knowledge graphs are essential tools for representing real-world facts and finding wide applications in various domains. However, the process of constructing knowledge graphs often introduces noises and errors, which can negatively impact the performance of downstream applications. Current methods for knowledge graph error detection primarily focus on graph structure and overlook the importance of textual information in error detection. Therefore, this paper proposes a novel error detection framework that combines both structural and textual information. The framework utilizes a confidence module for error detection while generating knowledge embeddings. The performance of this approach outperforms baseline methods in error detection and link prediction experiments, particularly achieving state-of-the-art performance in the error detection task.
Graph Neural Networks (GNNs) are vulnerable to backdoor attacks, where adversaries implant malicious triggers to manipulate model predictions. Existing graph backdoor attacks are susceptible to defense mechanisms or robust classifiers because they rely on subgraph injection or structural perturbations, e.g., creating additional edges to attach backdoor triggers to the original graph. To enhance the stealthiness of graph backdoors, we propose SPEAR, a novel structure-preserving graph backdoor attack that avoids modifying the graph’s topology. SPEAR operates within a limited attack budget by selectively perturbing node attributes while ensuring the triggers exert significant influence through a global importance-driven feature selection strategy. Additionally, a neighborhood-aware trigger generator is employed to underpin a high attack success rate by utilizing semantic information from the neighborhood. SPEAR amplifies effectiveness and stealthiness by combining subtle yet impactful attribute manipulation with a refined trigger generation mechanism. Extensive experiments demonstrate that SPEAR achieves state-of-the-art effectiveness in bypassing defenses on real-world datasets, establishing it as a potent and stealthy backdoor attack for graph-based tasks.
Graph Neural Networks (GNNs) are vulnerable to perturbations in both edges and attributes by fraudsters attempting to evade detection. A low-cost and effective perturbation strategy involves establishing connections with benign users and providing as little information as possible, leading to a graph with noisy structure and absent attributes. We formulate a novel problem as learning in Graphs with Noisy structures and Absent node attributes (LGNA), for which no existing methods are specifically designed. To mitigate this gap, we propose a reliable graph learning framework called RENA, which implements a “Dilution of Unreliable Information” approach for the LGNA task. The core principle of RENA is to utilize more reliable information to decrease the proportion of unreliable information, thus diluting its impact. Specifically, only the observed node attributes and unconnected node pairs are considered reliable, while imputed attributes and connected node pairs are deemed unreliable. We first randomly sample a large number of unconnected node pairs and fewer connected pairs to create different structural views to supervise structure learning and dilute the impact of noisy edges. Next, we apply a graph autoencoder framework, assigning higher weights to the observed attributes and lower weights to the imputed attributes during the reconstruction process, thereby diluting the impact of imputation noise. Experiments show that our method outperforms state-of-the-art baselines on LGNA scenarios and conventional incomplete graph learning tasks. Code is available at https://github.com/lxx01110/RENA.
Anaerobic methanotrophic (ANME) microbes play a crucial role in the bioprocess of anaerobic oxidation of methane (AOM). However, due to their unculturable status, their diversity is poorly understood. In this study, we established a microfluidics-based epicPCR (Emulsion, Paired Isolation, and Concatenation PCR) to fuse the 16S rRNA gene and mcrA gene to reveal the diversity of ANME microbes (mcrA gene hosts) in three sampling push-cores from the marine cold seep. A total of 3725 16S amplicon sequence variants (ASVs) of the mcrA gene hosts were detected, and classified into 78 genera across 23 phyla. Across all samples, the dominant phyla with high relative abundance (>10%) were the well-known Euryarchaeota, and some bacterial phyla such as Campylobacterota, Proteobacteria, and Chloroflexi; however, the specificity of these associations was not verified. In addition, the compositions of the mcrA gene hosts were significantly different in different layers, where the archaeal hosts increased with the depths of sediments, indicating the carriers of AOM were divergent in depth. Furthermore, the consensus phylogenetic trees of the mcrA gene and the 16S rRNA gene showed congruence in archaea not in bacteria, suggesting the horizontal transfer of the mcrA gene may occur among host members. Finally, some bacterial metagenomes were found to contain the mcrA gene as well as other genes that encode enzymes in the AOM pathway, which prospectively propose the existence of ANME bacteria. This study describes improvements for a potential method for studying the diversity of uncultured functional microbes and broadens our understanding of the diversity of ANMEs.