Dioscorea alata L., the most widely cultivated yam species, exhibits dioecy (XY system) and a strong male-biased sex ratio. These two major constraints limit parental combinations and hinder breeding progress. To get insight into the sex chromosome structure and evolution in D. alata, we present a haplotype-resolved, near-complete genome assembly of the male cultivar "Kabusa," spanning 958 Mb across 40 chromosomes. DalaChr6A is identified as the Y chromosome which exhibits early signs of heteromorphism, including a slight reduction in size (∼400 Kb difference) and a potential centromere shift relative to the X chromosome. Our findings also reveal a sex chromosome turnover between D. alata and D. rotundata. The sex chromosomes of D. alata are evolutionarily young (approximately 4.32 million years ago) and emerged after the divergence from the D. rotundata lineage. The sex-determining region (SDR) is refined to approximately 7.6 Mb, representing approximately 44% of the Y chromosome. It contains several inversions and a divergence gradient was observed across these inversions. INV4 identified as the oldest pericentric inversion likely marks the early step in D. alata SDR evolution. Despite structural divergence, both X and Y chromosomes remain transcriptionally active. Among the 231 genes annotated in the SDR, 97 are sex-biased and enriched in functions associated with floral organ formation and hormonal signaling pathways. This study enhances our understanding of sex chromosome evolution and sex determination in dioecious plants. It provides a gold-standard reference genome for the Dioscorea genus and lays the foundation for accelerated breeding and genetic improvement in yam.
The strategy for Energy Efficient Ethernet determines when to enter and leave the power-saving mode, thereby directly impacting both the energy savings and the incurred latency of frames. However, due to the strong dependency of the strategy performance on the network traffic, existing strategies for Energy Efficient Ethernet need either (i) appropriately static parameter configuration under certain traffic loads, or (ii) parameter adaptation mechanisms based on the traffic prediction models that assume specific traffic distribution. Consequently, existing strategies hardly maintain consistent high performance under variable traffic in reality. To address this issue, we incorporate reinforcement learning into the design of the Energy Efficient Ethernet strategy and propose the reinforcement learning based periodic strategy (RLPS). Specifically, RLPS operates periodically with a dynamically adjusted cycle length. In each cycle, RLPS first transmits all buffered frames, then enters a selected power-saving mode for the remaining duration. Rather than directly outputting power-saving mode transition decisions, RLPS learns the time length of each cycle online for effectively capturing the impacts of traffic variation. This periodic approach enables power consumption to be optimized within each cycle using learned information while also reducing the overhead of online learning. Extensive simulations driven by synthetic traffic and real traces show that RLPS outperforms existing strategies, reducing the power consumption by up to similar to 60.8% while maintaining consistent high performance across diverse traffic loads and distributions.
The short video market is currently growing rapidly. As a typical bandwidth-intensive application, short video streaming can easily cause bandwidth bottlenecks in servers. Hence, we propose a prefetching strategy to improve the quality of experience (QoE) and reduce the wastage of bandwidth, utilizing computility for performance optimization. We reveal that the state-of-the-art Dashlet prefetching strategy fails to maximize the weighted sum of QoE and bandwidth usage (i.e., utility). To overcome this limitation, the proposed comprehensive multistep prefetching strategy (CMPS) for short video streaming computes the prefetching urgency of chunks and combines it with expected utility to comprehensively evaluate multistep decision sequences, achieving higher QoE and bandwidth utilization than Dashlet. The extensive evaluations show that the proposed CMPS improves the average utility across diverse scenarios by 16.1% compared with Dashlet.
MOTIVATION:Accurately estimating changes in binding free energy (ΔΔG) is critical for understanding protein-protein interactions (PPIs) and guiding rational protein design. Recent deep learning methods have achieved notable progress by pre-training on large-scale structural data. While some approaches explore structural flexibility through energy-based sampling or generative modeling, these strategies typically involve substantial computational cost and overlook data efficacy in terms of training data organization. RESULTS:We present USP-ddG, a unified structural paradigm for ΔΔG prediction built on a dual-channel architecture. The model incorporates three complementary components: (i) an inverse folding-based log-odds ratio, (ii) the empirical force field FoldX capturing side-chain packing energetics, and (iii) a geometric encoder that leverages Gaussian coordinate perturbation as a regularization strategy to improve robustness. To enhance representation capacity, we introduce a framework that integrates feed-forward network (FFN) and Mixture-of-Experts (MoE) to model domain-invariant and -specific features, respectively. We further propose CATH-guided Folding Ordering (CFO), a data efficacy strategy that organizes samples to mitigate catastrophic forgetting and data distribution bias. USP-ddG consistently outperforms existing state-of-the-art methods on the SKEMPI v2.0 benchmark, including the challenging hold-out CATH test set. It achieves superior accuracy on both single- and multi-point mutations and demonstrates strong performance in antibody affinity optimization against H1N1 and HER2, and in predicting the impact of SARS-CoV-2 variants on hACE2 binding. Ablation studies confirm the contribution of each component. These results highlight USP-ddG as a robust and data-efficient framework for modeling mutational effects on PPIs. AVAILABILITY AND IMPLEMENTATION:USP-ddG is available at https://github.com/ak422/USP-ddG.
Early diagnosis of arrhythmia, a common cardiovascular condition, is crucial for improving prognosis. Electrocardiogram (ECG) is widely used as a non-invasive diagnostic tool. However, Computer-Aided Diagnosis of rare arrhythmias faces significant challenges due to the severe scarcity of samples for these rare disease classes. To tackle this, we propose a Mamba-based Prototypical Contrastive Learning framework, which can simultaneously identify both common and rare classes under the setting of generalized Few-Shot Learning (FSL). It primarily consists of: (1) the Mamba-based Spatio-Temporal Feature Fusion Network (MST), which integrates spatial features from multi-scale convolutions and temporal dynamics from bidirectional Mamba for ECG modeling; (2) the Prototypical Contrastive Learning framework with Augmented Feature Separation (PCAS), which employs a prototype augmentation strategy with an Augmented Prototype Consistency Loss to optimize prototype representations, and an Separation-Tuned Contrastive Loss to enhance intra-class compactness and inter-class distinctnessy, mitigating the risk of class collapse. Extensive experiments on publicly available datasets PTBXL and Chapman demonstrate the effectiveness of MST-PCAS, achieving superior rare-class recognition accuracies of 79.13% and 50.72%, respectively, for ECG arrhythmia classification.
Accurate glioma grading from Magnetic Resonance Imaging (MRI) is critical for early diagnosis and effective treatment planning. Existing Artificial Intelligence (AI) methods, such as Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs), struggle to capture edge details, overlook volumetric spatial context, and lack sensitivity to irregular shapes. Furthermore, inadequate self-attention mechanisms cause ViTs to fail to prioritize semantically meaningful tokens. These limitations can hinder both diagnostic accuracy and generalization. To address these challenges, we propose the Edge-Aware Transformer (EA-Trans) for glioma grading. Our approach integrates four key modules. First, the Edge Sensitive Tokenization (EST) module employs edge enhancement filters using depthwise CNNs to capture edge information. Second, the Shared Axis Spatial Features Alignment (SASFA) module processes MRI volumes along the depth, height, and width axes, preserving spatial consistency and context. Third, the Volumetric Multi-Scale Wavelet Transform Convolution (VMWTC) module employs Wavelet Transform Convolutions (WT-CNNs) to extract shape-sensitive multi-scale features. Fourth, the Adaptive Self-Attention (ASA) module integrates an Interquartile Range Token Selection (IQR-TS) strategy to focus on semantically relevant tokens. Experiments on the publicly available Brain Tumor Segmentation (BraTS2020) and University of California San Francisco Preoperative Diffuse Glioma MRI (UCSF-PDGM) datasets demonstrate that the proposed model achieves accuracies of 95.7% and 95.2%, respectively, outperforming 12 baseline methods. Notably, the model, with only 3.176 million parameters and 1.970 billion floating-point operations, enhances glioma grading accuracy, supporting AI-driven neuroimaging diagnostics in medical engineering. Interpretability analysis and application to Alzheimer’s disease classification further suggest its potential applicability to other neuroimaging tasks.
Parkinson’s disease (PD) is a progressive degenerative neurological disease in which timely diagnosis is essential for slowing clinical deterioration. Existing multimodal neuroimaging methods based on structural MRI (sMRI) and functional MRI (fMRI) often learn suboptimal modality-specific representations and rely on shallow fusion strategies that underutilize cross-modal dependencies. To solve these challenges, a Spatial-Temporal dual-pathway network with multi-scale attention Fusion is proposed for PD diagnosis (STFusion). In particular, a hybrid CNN-Transformer branch models both local and global structural patterns from sMRI, while a spatial-temporal Transformer branch characterizes dynamic functional connectivity from fMRI. A cross-modality multi-scale attention integration (CMAI) block is further introduced to adaptively integrate complementary information across modalities. Comprehensive evaluations on a public PPMI cohort and an external cohort show that STFusion achieves the best performance, reaching accuracies of 0.926 and 0.858, respectively.
Deep learning has achieved remarkable success in automated electrocardiogram (ECG) diagnosis. However, the long-tailed distribution of real-world ECG data remains a critical challenge. Existing methods mainly operate at the signal feature level and fail to exploit semantic relationships among diagnostic labels, which limits their ability to effectively recognize rare categories. To address this issue, we propose a Semantics-Guided Representation Learning (SGRL) framework, designed to leverage semantic priors to guide feature learning and thereby correct feature bias under long-tailed distributions. Specifically, we propose a Label Semantic Prompt Adaptation (LSPA) strategy, which introduces ECG-calibrated learnable parameters into label prompts to generate adaptive label semantic embeddings under a semantic-consistency constraint, thereby overcoming the limitations of coarse-grained label prompts in characterizing fine-grained intra-class morphological variations. Subsequently, we propose Semantic-Guided Representation Decorrelation (SGRD) that fuses ECG embeddings with adaptive label semantics and suppresses inter-class semantic correlations to obtain discriminative class semantic anchors, thereby improving tail-class separability. Extensive experiments on multiple public datasets demonstrate the significant superiority of SGRL in long-tailed ECG classification tasks. Our code is available on https://github.com/gfywudi/SGRL.
Visual Floorplan Localization (FLoc) struggles with severe structural aliasing caused by repetitive minimalist layouts. This occurs because physically distant poses share highly similar visual-geometric features, which degrades spatial separability and angular discriminability. While existing methods attempt to mitigate these ambiguities by relying on costly semantic annotations, the resulting performance gains remain inherently limited. To address the above issues, we propose DisCo-FLoc, a semantic-free method for visual-geometric Contrastive Disambiguation. First, we introduce a depth-aware Ray Regression Predictor (RRP) that serves as a dense-to-ray geometric projector. By explicitly suppressing visual clutter along the vertical dimension, RRP projects monocular RGB images into 2D ray primitives, which are matched with floorplans to produce geometry-aware FLoc candidates. Second, to resolve the remaining ambiguity among these candidates, we propose a spatially perturbed contrastive objective to align RGB images with local floorplan structures and formulate a visual-geometric compatibility function. In particular, we meticulously construct positive and negative samples at both positional and directional levels through SE(2) pose perturbations for contrastive learning, effectively achieving pose smoothness, spatial separability, and angular discriminability. The compatibility function enables DisCo-FLoc to disambiguate FLoc by using richer visual context beyond pure geometric layouts, without requiring any semantic annotations. Extensive experiments on two challenging visual FLoc benchmarks demonstrate that DisCo-FLoc significantly outperforms state-of-the-art semantic-based methods, especially narrowing the performance gap between positional and directional FLoc accuracy.
With the growing ubiquity of video streaming, ensuring a fair and high quality of experience (QoE) for users has emerged as a shared concern among video content providers. State-of-the-art video delivery systems achieve QoE fairness through bottleneck bandwidth allocation across multiple video streams, all based on the assumption of a unified congestion control (CC) protocol. However, the widespread use of heterogeneous CC protocols on the Internet not only disrupts QoE fairness among video streams but also poses challenges in achieving fast convergence under dynamic bandwidth. To address these issues, we propose a QoE-Fairness aware bandwidth allocation mechanism called Fabam, which establishes a unified QoE control plane across heterogeneous CC protocols. Fabam constructs independent virtual targets based on the real-time QoE of each video stream to achieve QoE fairness, and offers rapid convergence for the underlying CC protocols to improve efficiency. In addition, we propose a Deep Neural Network (DNN)-based multi-step mapping model aimed at balancing the performance and overhead of Fabam, thereby enhancing its deployment potential in practical applications. We implement Fabam on QUIC and integrate it with Dash.js. The evaluation results demonstrate the significant superiority of Fabam over the state-of-the-art approaches, including an enhancement of 44.01% in QoE fairness and an improvement of 36.39% in QoE efficiency. Meanwhile, Fabam-DNN maintains satisfactory QoE fairness while supporting multiple users at a low cost.
Drug repurposing represents a cost-effective strategy to identify novel therapeutic applications for existing pharmaceuticals, circumventing the protracted timelines of traditional drug discovery. While knowledge graph (KG) based methods excel at integrating heterogeneous biomedical data, they often struggle to harmonize high-level domain knowledge with fine-grained molecular mechanisms. We propose KGDDA, a multimodal framework designed for drug-disease association prediction that synergistically integrates KGs with medical ontologies. By leveraging an attention-driven fusion mechanism, KGDDA dynamically merges contextual topological embeddings with ontology-derived priors, enabling the adaptive capture of intricate drug-disease interactions. Extensive evaluations on two benchmark datasets demonstrate that KGDDA consistently outperforms state-of-the-art baselines in both predictive accuracy and generalization. Furthermore, case studies on head and neck cancer and small cell lung cancer validate KGDDA's ability to provide actionable mechanistic insights, highlighting its potential to accelerate therapeutic discovery and precision medicine.
In this paper, we study coreset construction for LASSO regression, where a coreset is a small, weighted subset of the data that approximates the original problem with provable guarantees. For unregularized regression problems, sensitivity sampling is a successful and widely applied technique for constructing coresets. However, extending these methods to LASSO typically requires coreset size to scale with O(\mathcal{G}d), where d is the VC dimension and \mathcal{G} is the total sensitivity, following existing generalization bounds. A key challenge in improving upon this general bound lies in the difficulty of capturing the sparse and localized structure of the function space induced by the \ell_1 penalty in LASSO objective. To address this, we first provide an empirical process-based method of sensitivity sampling for LASSO, localizing the procedure by decomposing the functional space into independent spaces, which leads to tighter estimation error. By carefully leveraging the geometric properties of these localized spaces, we establish tight empirical process bounds on the required coreset size. These techniques enable us to achieve a coreset of size \tilde{O}(\epsilon^{-2}d\cdot((\log d)^3\cdot\min(1,\log d/\lambda^2)+\log(1/\delta))), which ensures a (1\pm\epsilon)-approximation for any \epsilon,\delta\in(0,1) and \lambda > 0. Furthermore, we give a lower bound showing that any algorithm achieving a (1+\epsilon)-approximation must select at least \Omega(\frac{d\log{d}}{\epsilon^2}) rows in the regime where \lambda=O(d^{-1/2}). Empirical experiments show that our proposed algorithm is at least 4 times faster than the existing LASSO solver and more than 9 times faster on half of the datasets, while ensuring high solution quality and sparsity.
RNA-binding proteins (RBPs) regulate gene expression through dynamic and static interactions modulated by cellular environments. Traditional methods focus on either static or dynamic mechanisms in isolation, overlooking the inherent characteristics of protein-RNA interactions (PRIs). Here, we present DyStaBind, a multi-view framework capturing RNA’s dynamic sequence embeddings, dynamic structural characterization, and static conformational features. Within each view, parallel CNNs with varying kernel sizes extract multi-scale latent motifs, and cross-attention modules based on multilingual alignment complete the correlation among features at different levels. A pyramidal context-aware classifier with varying kernel sizes, residual connections, and max-pooling to capture long-distance dependencies for PRIs prediction. Benchmark evaluations across 261 eCLIP and CLIP-seq datasets confirm DyStaBind’s superior performance over state-of-the-art models, achieving robust generalization in both static cellular detection and dynamic cellular detection. Ablation experiments verify that the multilingual alignment fusion module effectively separates the features of positive and negative samples. The visualization analysis of the attention weights shows that the focus area of the model is highly consistent with the statistically significant functional motifs, confirming the reliability at the level of biological semantic analysis. In addition, the dynamic changes of sequence binding patterns in different cellular environments are accurately captured.
The existing contrastive learning methods widely adopt one-hot instance discrimination as pretext task for self-supervised learning, which inevitably neglects rich inter-instance similarities among natural images, then leading to potential representation degeneration. In this paper, we propose a novel image mix method, PatchMix, for contrastive learning in Vision Transformer (ViT), to model inter-instance similarities among images. Following the nature of ViT, we randomly mix multiple images from mini-batch in patch level to construct mixed image patch sequences for ViT. Compared to the existing sample mix methods, our PatchMix can flexibly and efficiently mix more than two images and simulate more complicated similarity relations among natural images. In this manner, our contrastive framework can significantly reduce the gap between contrastive objective and ground truth in reality. Experimental results demonstrate that our proposed method significantly outperforms the previous state-of-the-art on both ImageNet-1K and CIFAR datasets, e.g., 3.0% linear accuracy improvement on ImageNet-1K and 8.7% kNN accuracy improvement on CIFAR100. Moreover, our method achieves the leading transfer performance on downstream tasks, object detection and instance segmentation on COCO dataset. The code is available at https://github.com/visresearch/patchmix
Many deep learning methods represented by graph-based approaches achieve significant progress in drug repositioning. However, these graph-based methods face a critical limitation: they often fail in cold-start scenarios because the graph structure relies heavily on known association information from both the drug and disease sides. To address this challenge, we propose a bidirectional behavior learning strategy for drug repositioning, BiBLDR, an innovative framework that reformulates drug repositioning as a behavior sequence learning task. First, we construct bidirectional behavioral sequences based on drug and disease sides. Bidirectional behavior sequences ensure sufficient information for model learning in both drug and disease cold-start scenarios, while providing more precise feature representations for association prediction tasks. Subsequently, we propose a two-stage strategy for drug repositioning. In the first stage, we construct prototype spaces to characterise the representational attributes of drugs and diseases. In the second stage, these refined prototypes and bidirectional behavior sequence data are leveraged to predict potential drug-disease associations. This design allows BiBLDR to more robustly capture hidden pharmacological relationships from bidirectional behavioral sequences, delivering significant benefits in cold-start scenarios. Extensive experiments demonstrate that our method achieves state-of-the-art performance on benchmark datasets. Meanwhile, BiBLDR demonstrates significantly superior performance compared to previous methods in cold-start scenarios.
Nucleic acid-binding proteins (NABPs), including single-stranded DNA-binding proteins (SSBs), double-stranded DNA-binding proteins (DSBs), and RNA-binding proteins (RBPs), play pivotal roles in essential biological processes. Many computational approaches have been developed for the classification and identification of these NABP subtypes through machine learning and deep learning techniques. However, these methods exhibit limitations in accurately classifying NABP types, particularly in cross-prediction scenarios such as misclassifying SSBs as RBPs and vice versa. To address these challenges and improve the prediction accuracy of different types of NABPs, we proposed AttentionNABP, a deep learning framework that integrates convolutional neural networks, a multi-head self-attention mechanism, and long short-term memory networks. Our method offers two complementary prediction strategies: a hierarchical cascaded classifier and a simultaneous multi-class classifier. The hierarchical approach achieves prediction accuracies of up to 96.39 ± 0.44
The growing availability of multimodal Electronic Health Records (EHRs) data provides promising research on multimodal representation learning. Leveraging multimodal EHRs data and extracting inter-modality correlation among them has achieved acceptable performance for the mortality risk prediction task. However, the existing methods do not address the problem of the imbalanced convergence for each modality encoder. Furthermore, the inter-modality correlation fusion among different modalities is limited. This work proposes multimodal representation learning with pre-training and inter-modality correlation-aware fusion to address the imbalanced training convergence and capture the inter-modality correlation among all modalities’ representations. Specifically, the proposed method comprises two main stages. In the first stage, we pre-train the time series and demographic information modalities’ encoders separately and fine-tune the pre-trained Biomedical Bidirectional Encoder Representations Transformer (BioClinicalBERT) to provide a balanced training convergence. In the second stage, the pre-trained encoders from the first proposed stage are utilized to encode the time series, demographic, and clinical notes modalities. Then, we fuse their representations to extract the intercorrelation utilizing the low-rank multimodal fusion and the concatenation operation to provide the final patient representation. We evaluate the proposed framework on the MIMIC-III dataset. The experimental results demonstrate that the proposed framework outperforms the state-of-the-art models for mortality risk prediction.