With the development of foundational models, model compression has become a critical requirement. Various model compression approaches have been proposed such as low-rank decomposition, pruning, quantization, ergodic dynamic systems, and knowledge distillation, which are based on different heuristics. To elevate the field from fragmentation to a principled discipline, we construct a unifying mathematical framework for model compression grounded in measure theory. We further demonstrate that each model compression technique is mathematically equivalent to a neural network subject to a regularization. Building upon this mathematical and structural equivalence, we propose an experimentally-verified data-free model compression framework, termed Big2Small, which translates Implicit Neural Representations (INRs) from data domain to the domain of network parameters. Big2Small trains compact INRs to encode the weights of larger models and reconstruct the weights during inference. To enhance reconstruction fidelity, we introduce Outlier-Aware Preprocessing to handle extreme weight values and a Frequency-Aware Loss function to preserve high-frequency details. Experiments on image classification and segmentation demonstrate that Big2Small achieves competitive accuracy and compression ratios compared to state-of-the-art baselines.
While Multimodal Large Language Models (MLLMs) excel in semantic tasks, they frequently lack the "spatial sense" essential for sophisticated geometric reasoning. Current models typically suffer from exorbitant modality-alignment costs and deficiency in fine-grained structural modeling precision.We introduce SSR, a framework designed for Structured Scene Reasoning that seamlessly integrates 2D and 3D representations via a lightweight alignment mechanism. To minimize training overhead, our framework anchors 3D geometric features to the large language model's pre-aligned 2D visual semantics through cross-modal addition and token interleaving, effectively obviating the necessity for large-scale alignment pre-training. To underpin complex spatial reasoning, we propose a novel scene graph generation pipeline that represents global layouts as a chain of independent local triplets defined by relative coordinates. This is complemented by an incremental generation algorithm, enabling the model to construct "language-model-friendly" structural scaffolds for complex environments. Furthermore, we extend these capabilities to global-scale 3D global grounding task, achieving absolute metric precision across heterogeneous data sources. At a 7B parameter scale, SSR achieves state-of-the-art performance on multiple spatial intelligence benchmarks, notably scoring 73.9 on VSI-Bench. Our approach significantly outperforms much larger models, demonstrating that efficient feature alignment and structured scene reasoning are the cornerstones of authentic spatial intelligence.
Depression represents the predominant cause of disability worldwide and substantially influences patients’ affective vocal characteristics and emotional expressivity. The prediction of depressive states through auditory emotional cues has been extensively investigated in previous literature. However, the analysis of depressive conditions based on audio signals remains considerably challenging, and the transferability of emotion analysis frameworks to depression detection in existing studies is notably constrained. To mitigate this limitation, we introduce a robust semi-supervised learning framework for depression assessment utilizing audio data. This framework offers a novel perspective on the interplay between depressive states and emotion-related variations in speech. It comprises three integral components: pseudo-label generation, self-training based on pseudo-labels, and a depression classification mechanism that incorporates both local and global contextual information. The first module involves segmenting the audio stream into multiple clips and assigning pseudo-labels indicative of emotional tendencies. The second component constructs an emotion classifier to predict emotional propensity for each segment. The third component employs a transformer encoder to develop an attention generation module, effectively integrating global features to refine local predictions for the depression classification task. Empirical evaluations conducted on the DAIC-WOZ and EATD datasets demonstrate the superior performance of the proposed framework over existing classification approaches. Additionally, we conducted a comprehensive investigation into the impact of various feature representations on model efficacy. The proposed framework provides an effective and analytically rigorous solution for automated depression classification.
Recently, Neural Ordinary Differential Equations (NODEs) have emerged as a promising paradigm for lightweight neural networks. However, this approach remains constrained by the numerical instability of ODE solvers, particularly when applied to deep networks. In this paper, we propose GhostODE, a novel lightweight network that integrates Heun's method with Ghost modules. First, we propose a direct differentiable Heun's algorithm to improve numerical stability and accuracy with minimal computational overhead. Second, we employ Ghost modules to optimize pointwise convolutions, effectively alleviating the parameter bottleneck. Finally, Knowledge Distillation (KD) is applied to further enhance GhostODE's overall performance. Extensive experiments across diverse tasks, including image classification, object detection, semantic segmentation, and real-world edge deployment, validate the efficacy of the proposed GhostODE. Furthermore, the integration of Heun’s method and Ghost modules enables GhostODE to maintain high accuracy while reducing computational complexity, making it suitable for real‐time applications on mobile devices. Notably, GhostODE achieves only 0.45M parameters (86.5% reduction vs. MobileNetV1) and reaches 127.2FPS on Android platforms, demonstrating its practicality for edge intelligence. Code is available at https://github.com/01LuGuang/GhostODE.
U-like networks have become fundamental frameworks in medical image segmentation through skip connections that bridge high-level semantics and low-level spatial details. Despite their success, conventional skip connections exhibit two key limitations: inter-feature constraints and intra-feature constraints. The inter-feature constraint refers to the static nature of feature fusion in traditional skip connections, where information is transmitted along fixed pathways regardless of feature content. The intra-feature constraint arises from the insufficient modeling of multi-scale feature interactions, thereby hindering the effective aggregation of global contextual information. To overcome these limitations, we propose a novel Dynamic Skip Connection (DSC) block that fundamentally enhances cross-layer connectivity through adaptive mechanisms. The DSC block integrates two complementary components: (1) Test-Time Training (TTT) module: This module addresses the inter-feature constraint by enabling dynamic adaptation of hidden representations during inference, facilitating content-aware feature refinement. (2) Dynamic Multi-Scale Kernel (DMSK) module: To mitigate the intra-feature constraint, this module adaptively selects kernel sizes based on global contextual cues, enhancing the network’s capacity for multi-scale feature integration. The DSC block is architecture-agnostic and can be seamlessly incorporated into existing U-like network structures. Extensive experiments demonstrate the plug-and-play effectiveness of the proposed DSC block across CNN-based, Transformer-based, hybrid CNN-Transformer, and Mamba-based U-like networks. The code is available at https://github.com/BlackJack-Cao/U-like-Networks-with-DSC.
Deploying deep neural networks with massive parameter counts on resource-constrained edge devices remains a significant challenge due to limited computational power, storage capacity, and energy efficiency. Knowledge distillation has emerged as a promising model compression technique, enabling the transfer of knowledge from a large teacher model to a compact student model. Traditional relation-based knowledge distillation methods focus on distilling inter-instance relational knowledge—typically, pairwise similarities between samples—but often overlook the asymmetry in each sample’s contribution to these relations. In reality, the relational information derived from a pair of samples may hold different importance for each sample within the feature space. To address this issue, we propose Weighted Sample Correlation Knowledge Distillation (WSCKD), a novel approach that explicitly models the asymmetric contributions of samples to relational knowledge. WSCKD leverages pairwise similarities in the source (teacher) embedding space as transferable knowledge and introduces two asymmetric loss functions: the unidirectional weighted distillation loss and the bidirectional weighted distillation loss. These losses enable the student model to prioritize more informative relationships during training, without imposing constraints on the manifold structure of the student’s embedding space. Extensive experiments and ablation studies on multiple visual image recognition benchmarks demonstrate that WSCKD outperforms sixteen distillation methods.
The solubility of active pharmaceutical ingredients is vital throughout the drug design, development processes and manufacture. However, solubility prediction remains a challenging task in the pharmaceutical field. Therefore, BCS class II drugs solubility prediction model was developed on the basis of the machine learning algorithms and molecular descriptors through Bayesian Optimization, cosine similarity and sparse principal component analyses, revealing XGBoost model exhibited the better accuracy and suitability. Besides, the generalization of XGBoost model was confirmed by the solubility data prediction in the uncommon solvents and unseen solutes. Influences of molecular descriptors on the predicted solubility data were evaluated through Shapley Additive Explanations analysis, exposing the temperature exhibited a positive effect on the predicted solubility and the double bonds number of the solvent molecule presented a negative effect on the predicted solubility data. The various molecular descriptor contributions to the solubility prediction of XGBoost model were analyzed through feature importance, exposing the molecular descriptor contributions followed the order: Chi0 > SMR_VSA1 > MolMR > ExactMolWt > T > NumValenceElectrons > fr_C_O. In addition, it revealed the studied molecular descriptors must synergistically contribute to the solubility data prediction of XGBoost model according to prediction results comparison of simple and original XGBoost models.
Understanding the neural mechanisms underlying biological motion perception remains a significant challenge in neuroscience. To further explore this mechanism, we construct the BioMotionNet model using real bio-neural data from the MT to MST regions in macaques. To characterize neuron activity within particular time windows, we propose the window learning strategy, which employs windowed learning to extract crucial information related to specific events or stimuli. By analyzing the connectivity structure of the BioMotionNet model, we identify regular projection patterns from MT to MST, reflected in the varying response characteristics of MT neurons based on their projection strength to different MST neuron populations. Our data/codes are available at https://github.com/BrainCogLab/MT_MST.
The Transformer architecture has achieved remarkable success in computer vision and natural language processing. However, its application to time series forecasting frequently results in performance degradation, occasionally underperforming even simple linear models. Prior analyses have predominantly attributed this limitation to suboptimal embedding designs that fail to construct a well-structured latent space capable of effectively capturing intricate temporal dependencies. Most existing architectures rely on linear embeddings for input mapping, but these transformations often fail to project raw time series data onto a high-dimensional manifold, resulting in latent representations that poorly capture complex temporal structures and lead to attention degeneration. To overcome these challenges, we propose Structured Latent Projection (SLP), an enhanced embedding method that generates a rich, structured latent space from raw time series data. By mapping input sequences to a high-dimensional manifold that captures multi-scale temporal dependencies and inter-variable interactions, SLP improves the efficiency and robustness of the self-attention mechanism. We integrate SLP into the Transformer architecture, resulting in a model called LatentBridge. Extensive experiments on 13 real-world datasets show that LatentBridge consistently achieves state-of-the-art performance in both long-term and short-term forecasting.
Event co-occurrences have been proven effective for event argument extraction (EAE) in previous studies; however, few have considered intra- and inter-event role correlations. Since role varies among different event types, event structure heterogeneity and overlap pose significant challenges to EAE. To address this issue, we propose a Role Correlation Structure-Enhanced model for Multi-Event Argument Extraction (RoSE), capable of capturing both heterogeneity and overlap of event structures through modeling role correlations. The proposed RoSE model employs a joint context-prompts input, role-centric graph-guided encoder (RoGE), and role-specific information fusion (RoIF). The RoGE is designed to enhance the intra- and inter-event role correlation between prompts and their corresponding event contexts. The RoIF module utilizes intra-event role information to improve multi-event arguments extraction. Extensive experiments on four widely-used benchmarks (RAMS, WikiEvents, MLEE, and ACE05) demonstrate that our proposed approach achieves state-of-the-art performance, validating the effectiveness of incorporating both intra- and inter-event role correlations.
Micro-expression recognition (MER) reveals genuine emotions and is widely applied in fields such as depression detection and clinical psychological assessment. However, due to the short duration, low intensity of muscle movements, and small affected facial regions of micro-expressions (MEs), MER presents significant challenges. Despite recent advancements in neural network that have improved the performance of MER, existing network models still exhibit limitations owing to issues such as the small-scale of MEs datasets, insufficient training data, and class imbalance. In MER, the optical flow features reflect subtle changes in muscle movements between consecutive frames in video segments, the AUs are closely related to subtle changes in facial expressions, and the combinations of different AUs can reflect different facial expressions.We propose a novel framework called the AU-guided three-stream fusion network (AUTONet). The proposed framework constructs three subnetworks, one pre-trained on a large-scale image dataset and one shallow CNN, are used for extracting optical flow features, and another CNN based on the AUs intensity matrix is used for extracting AUs during MEs occurrences. The fusion module within the framework effectively utilizes the complementarity of multimodal features by creating a composite feature vector, thereby achieving better performance in MER. In addition, an AU-guided gaussian enhancement module guided by AUs is designed in the framework, which selectively enhances key regions in the optical flow using the position information of AUs. The experimental results indicate that AUTONet exhibits excellent performance on CASME II and SAMM datasets. Meanwhile, its low computational complexity and robustness to four types of perturbations show the practical viability in real-world conditions. Our work provides a promising direction for the combination of optical flow and AUs in MER.
Rectal cancer necessitates effective treatment strategies, with radiation therapy (RT) being crucial. Manual generation of intensity-modulated radiation therapy (IMRT) plans is time-consuming and expertise-dependent. This paper introduces an innovative approach for automatic IMRT planning, featuring three key elements: a singularity coding method, a cascaded model with neural memory Ordinary Differential Equation (nmODE), and a new evaluation criterion. The coding method reduces input data dimensions, representing 3D spatial information through 2D images and lowering computational cost. The cascaded model, with Dose Prediction and Fluence Map Prediction sub-models, integrates nmODE blocks to enhance nonlinearity. Ablation studies highlight the effectiveness of the cascaded structure and nmODE. Collaborative training strategies and a dual encoder in the Fluence Map Prediction Model (FPM) facilitate end-to-end learning. The new evaluation criterion calculates the error in the tumor target region. Experiments on in-house rectal cancer datasets show superior accuracy and efficiency, providing a new method for automated IMRT planning. Our source code is available at https://github.com/XiangjieTan25/nmODE-FluencePrediction.
This paper briefly studies the conditions for the coexistence of multiple attractors in recurrent neural networks with activation function RELU. In this paper we decompose the weight matrix into W+ and W-. Specifically, the entries in W+ come from the diagonal and positive off-diagonal elements of W, and the eigenvalues of matrix W+ play a crucial role in determining the existence and expressions of the corresponding attractor. By analyzing the eigenvalues of both matrix W+ and W, we obtain the coexistence conditions of different types of attractors in two-dimensional and three-dimensional spaces, as well as the conditions for determining the number of coexisting attractors in an n-dimensional network. Moreover, all research results were rigorously validated through simulation. Finally, we demonstrate the practical applications of coexisting continuous attractors for sequence prediction.
Test-Time Adaptation (TTA) enables models trained on a source domain to adapt online to unlabeled test data under distribution shifts. While recent TTA methods have moved beyond static settings and begun to consider continual domain shifts, they primarily address distribution drift and fail to account for class imbalance in dynamic scenarios. In real-world test-time streams, class imbalance and continual domain shifts often occur at the same time and interact with each other. In this paper, we propose a novel Balanced and Prototype-Guided Test-Time Adaptation (BP-TTA) method, which combines batch-balanced sampling with prototype-guided adaptation to handle the class imbalance and continual domain shift problems. BP-TTA constructs balanced adaptation batches by integrating current samples with high-confidence historical instances, effectively mitigating bias toward dominant classes and stabilizing online updates. Meanwhile, BP-TTA maintains evolving class prototypes during inference and leverages prototype similarity as a constraint for model adaptation, thereby improving the reliability of pseudo-labels and enhancing the stability of online updates under persistent domain shifts. Extensive experiments demonstrate that BP-TTA consistently outperforms state-of-the-art TTA methods in dynamic test-time streaming settings.
Session-based recommendation (SBR) in service computing is pivotal in predicting a user's next action based on their current anonymous session. While Graph Neural Network (GNN)-based methods have shown promise in capturing intricate item transformation relationships within sessions, they often fall short in accurately modeling user preferences. This is primarily due to the common practice of solely considering the last item in the session as the user's current interest, neglecting potentially valuable information embedded in other session items which is essential for capturing user global preferences. Moreover, existing models typically optimize performance solely through cross-entropy loss between predicted items and ground truth labels, while overlooking latent valuable knowledge embedded in intermediate features and item-item relationships that lends support to the model in accurately capturing and modeling user preferences. To address these shortcomings, we propose Multi-Scale Collaborative Distillation (MSCD) for SBR. Our approach introduces a current interest adaptive selection module, which dynamically selects appropriate item embeddings as session-local embeddings by evaluating the importance of each item within the session. This allows for a more accurate capture of the user's current true preferences. Additionally, we propose collaborative knowledge distillation, where multiple models are trained concurrently, enabling the transfer of three types of knowledge including response-based, feature-based, and relationship-based knowledge between models, thereby enriching the model's understanding of user preferences. Experimental evaluations conducted on three popular SBR datasets demonstrate that our MSCD model outperforms recent state-of-the-art methods in terms of recommendation accuracy.
Entity hallucination poses a major challenge in radiology report generation (RRG), particularly for 3D CT scans where complex spatial contexts amplify factual errors. To address this, medical entity phrases serve as key carriers for multi-modal prompting, integrating expert knowledge into the vision-language model. Current methods use unified cross-attention for volume-phrase alignment, failing to account for anatomical specificity during the alignment process. In this work, we introduce the Dual-stream Entity Alignment Reporting network (DEAR) that separately models organ and lesion entities to resolve anatomical bias. Specifically, the dual-stream entity aligner is designed to partition medical entity phrases into organ and lesion streams, feeding them into separate cross-attention blocks in parallel to achieve fine-grained volume–phrase alignment. For structurally regular and spatially stable organ entities, an organ-guided cross-attention (OGCA) block is proposed to enforce structural consistency by retrieving the top-k voxel tokens via volume–phrase similarity and preserving spatial connectivity through morphological dilation. Meanwhile, a lesion-guided cross-attention (LGCA) block is introduced for structurally irregular and spatially variable lesion entities, enhancing anomaly sensitivity through phrase-weighted attention and refining discriminative boundaries via 3D residual Laplacian filtering. Experiments demonstrate that DEAR significantly reduces entity hallucinations and improves clinical factuality in 3D RRG benchmarks.
The gradual deployment of Wi-Fi 7/8 multi-link operation (MLO) will lead to long-term coexistence between legacy non-MLO stations (STAs) and MLO-capable STAs in WLANs. This mixed deployment makes throughput optimization challenging because legacy STAs follow single-link contention and transmission, whereas MLO-capable STAs can exploit multiple links with richer access opportunities. Existing learning-based methods usually treat such networks as homogeneous systems and directly map the current observation to a complete MAC action, which cannot faithfully represent both legacy single-link and MLO multi-link behaviors. To address this issue, we propose EvoOMG, an evolution-oriented multi-agent guidance framework for heterogeneous legacy-and-MLO Wi-Fi networks. EvoOMG reformulates throughput optimization as a standard-constrained staged multi-agent decision problem. Each agent encodes recent channel, queue, contention, and transmission histories, first generates contention guidance, and then produces aggregation guidance conditioned on the preceding access stage and standard-specific feasibility constraints. This autoregressive design follows the Wi-Fi MAC order of “contention before transmission” while preserving distinct protocol behaviors of legacy and MLO-capable STAs. NS-3 evaluations show that EvoOMG improves scheduled goodput, convergence stability, and MLO link utilization over static enhanced distributed channel access (EDCA), one-step MADDPG, and independent-learning baselines, achieving substantial performance gains in representative mixed-standard scenarios.
Recent years have seen remarkable progress in deep learning on 3D point clouds, with hierarchical architectures becoming standard. Most work has focused on developing increasingly complex operators, such as self-attention, while enhancing the representational capacity of efficient point-wise MLP-based backbones has received less attention. We address this issue by proposing a differentiable module that learns to impose a task-driven canonical structure on local point sets. Our proposed SMA (Sort-Mix-Attend) layer dynamically serializes a neighborhood by generating a geometric basis and using a differentiable sorting mechanism. This enables an efficient MLP-based network to model rich feature interactions, adaptively modulating features prior to the final symmetric aggregation function. We demonstrate that SMA effectively enhances standard backbones for 3D classification and segmentation. Specifically, integrating SMA into PointNeXt-S achieves an Overall Accuracy (OA) of 88.3% on the challenging ScanObjectNN dataset, an improvement of 0.6% over the baseline. Furthermore, it boosts the classic PointNet++ architecture by a significant 5.2% in OA. We also introduce a highly efficient SMA-Tiny variant that achieves 86.0% OA with only 0.3 M parameters, proving the structural superiority, computational cost-effectiveness, and practical significance of our method for real-world 3D perception tasks.
Diffusion models generate high-quality images but pose serious risks like copyright violation and disinformation. Watermarking is a key defense for tracing and authenticating AI-generated content. However, existing methods rely on threshold-based detection, which only supports fuzzy matching and cannot recover structured watermark data bit-exactly, making them unsuitable for offline verification or applications requiring lossless metadata (e.g., licensing instructions). To address this problem, in this paper, we propose Gaussian Shannon, a watermarking framework that treats the diffusion process as a noisy communication channel and enables both robust tracing and exact bit recovery. Our method embeds watermarks in the initial Gaussian noise without fine-tuning or quality loss. We identify two types of channel interference, namely local bit flips and global stochastic distortions, and design a cascaded defense combining error-correcting codes and majority voting. This ensures reliable end-to-end transmission of semantic payloads. Experiments across three Stable Diffusion variants and seven perturbation types show that Gaussian Shannon achieves state-of-the-art bit-level accuracy while maintaining a high true positive rate, enabling trustworthy rights attribution in real-world deployment. The source code have been made available at: https://github.com/Rambo-Yi/Gaussian-Shannon
Diffusion probabilistic models have effectively addressed the ill-posed nature of cardiac magnetic resonance imaging (CMRI) super-resolution (SR) by learning high-resolution image distributions from low-resolution inputs. However, the iterative sampling process in these models often suffers from slow inference speeds, as well as limitations in the quality and structural consistency of the generated images. To address these challenges, we propose a continuous-time conditional diffusion model (CCDM) for blind CMRI SR. Specifically, we propose a continuous-time conditional diffusion module that reduces the time consumption of the diffusion probability model by maintaining the mean and variance of the data in the forward process. Meanwhile, we design a cascaded residual attention network as a feature extractor to enhance the model’s discriminative power and feature representation capabilities. To further elevate image fidelity, we propose an image quality loss module that integrates a score matching loss, significantly improving detail reconstruction and overall perceptual quality. Furthermore, we develop a hybrid score predictor that approximates the conditional score function via a hybrid parameterized denoising network, facilitating efficient CMRI generation through probability flow sampling. Extensive experimental results demonstrate that compared to existing diffusion model-based SR methods, our CCDM achieves significant improvements in SR quality while substantially reducing time consumption.