Transformer-based visual trackers have revolutionized the field by modeling robust long-range dependencies. However, they face a fundamental trade-off: global attention mechanisms suffer from quadratic computational complexity, while efficient window-based alternatives often restrict the receptive field, lacking the specific inductive bias required to localize dynamic targets undergoing large spatial displacements. To address this limitation, we introduce Spin Window Attention (SWA). Unlike conventional vision Transformers that primarily use window shifting for image continuity, SWA is specifically designed for tracking. It efficiently fuses template and search region features through a strategic cyclic rotation of query tokens. By employing this rotation, the model performs a spatial scan. The scanning enables the tracker to capture targets that move across windows, which in turn maintains linear computational complexity. Building on SWA, we propose SpinTrack, a Transformer-based tracker that achieves an optimal balance between accuracy and speed. To facilitate token rotation, an adaptive positional encoding scheme is developed to handle spatial permutations and varying resolutions – a design that further improves robustness. Extensive experiments show that SpinTrack achieves an optimal accuracy-speed trade-off on multiple benchmarks, attaining a highly competitive average overlap (AO) of 75.9% on GOT-10k while running at 75 FPS on a single RTX 4090 GPU.
Sequential recommendation (SR) methods encode dynamic preferences from a user’s historical data, usually constructed as a discriminant paradigm. Recent studies mainly reconstruct it into a learning-to-generate paradigm through the Classifier-Free Guided diffusion model and achieve remarkable results. This paradigm performs a forward noise addition and reverse reconstruction process in the embedding space of the target item. The generated results depend on the guidance signal used during the reconstruction process. Precise guidance reduces the deviation between the generated item and user preferences. However, this guidance signal is constructed from the embedding of user historical interaction sequences, which remains affected by data sparsity. This study proposes a novel dual contrastive learning method for sequential recommendation within the learning-to-generate paradigm, marking the first attempt to address the impact of data sparsity on the guidance signal during the reverse process. Building on this, we proposed a framework named Joint Diffusion Model and Dual Contrastive Learning for Sequential Recommendation (JCLRec). This method optimizes the guidance signal through contrastive learning, thereby enhancing the quality of the generated results. Specifically, this paper analyzes the optimization differences between the InfoNCE loss commonly used in SR tasks and the reconstruction loss used in the learning-to-generate paradigm. It proposes a feasible contrastive loss under this paradigm while maintaining the feature of not relying on negative sampling. In addition, to ensure the effectiveness of the contrastive learning strategy under this loss, this paper constructs two contrastive views. The first view is constructed through artificial rules to reduce the influence of data sparsity on the guidance signal. The second view is constructed through an optimized hard positive sampling strategy to reduce the semantic deviation effect brought by the previous operation on the guidance signal. Finally, this paper explores an effective combination of contrastive views and the learning-to-generate paradigm, proposing the JCLRec framework. The effectiveness of JCLRec is evaluated with existing methods on three challenging datasets, where the improvement rate under the HR@(5, 10, 20) indicator is from 5.2% to 18.6%, and NDCG@(5, 10, 20) is from 2.9% to 23.4%.
Malignant tumors impose a substantial burden on global health, with an urgent unmet need for effective targets to advance precision therapy. As evolutionarily conserved microtubule motor proteins, the kinesin family (KIF) orchestrates fundamental cellular processes (e.g., intracellular transport, cell division) and canonical signaling pathways, including Wnt and Hippo. Notably, recent studies have uncovered their emerging role in regulating tumor metabolism and reshaping the immune microenvironment, where dysregulated KIF expression drives aberrant tumor proliferation. Based on current research, here we synthesize the latest mechanistic insights into KIF-mediated tumor regulation and evaluate their translational potential as next-generation therapeutic targets and biomarkers for precision cancer therapy. KIF-targeted inhibitors have entered clinical trials across multiple cancer types, holding promise as novel anticancer agents to suppress tumor growth and progression, thereby providing valuable therapeutic options for clinical oncology practice.
Visual-language (VL) tracking leverages natural language descriptions to enhance visual object tracking. However, a critical yet overlooked issue is the inherent temporal heterogeneity between static language descriptions and dynamic visual sequences, which limits the efficacy of cross-modal fusion. To bridge this gap, we propose a novel Temporal-Align Visual-Language Tracker (TAVLT). Unlike prior works that simply concatenate multimodal tokens, TAVLT processes template, search, and language features separately to explicitly model their interrelations. Our framework introduces two core components: 1) a Temporal Semantic Enrichment Module (TSEM) that constructs a Spatio-Temporal Semantic Affinity Matrix to dynamically recalibrate language token weights, ensuring temporal alignment with the visual sequence; and 2) a lightweight Cross-Modal Feature Fusion Module (CMFFM) that leverages the affinity matrices for efficient feature integration. Furthermore, we introduce a Consistency Learning Balance (CLB) strategy to mitigate performance disparities across different perceptual constancy tasks (e.g., shape vs. color). Extensive experiments on LaSOT, LaSOT-ext, TNL2K, and OTB99-Lang demonstrate that TAVLT achieves state-of-the-art performance while maintaining high inference efficiency (49 FPS). Our approach effectively addresses the temporal misalignment in VL tracking, setting a new standard for robust and efficient cross-modal learning.
In recent years, the application of graph contrastive learning in recommendation systems has made significant progress. However, most current graph enhancement strategies still rely on manual experience design, lack flexibility and generalization ability, and are difficult to adapt to diverse recommendation tasks and complex graph structures. To this end, this paper proposes an adaptive multi-view contrastive recommendation algorithm (MAMGCL) based on a hybrid expert model. This method constructs an expert pool containing multiple enhancement strategies, introduces a two-way gating mechanism to achieve dynamic fusion of experts on the user side and the item side, further generates differentiated enhanced views, uses a multi-branch graph convolutional network to achieve node embedding expression, and finally combines sub-view-level contrastive learning for optimization. Experiments are carried out on three real-world datasets, verifying that the proposed method is superior to the existing mainstream contrastive recommendation models in terms of recommendation performance and robustness, demonstrating strong generalization ability and application potential.
Existing collaborative filtering recommendation paradigms primarily rely on co-occurrence matrices to mine user and item preference information. However, the limited interaction behavior in the real world making it difficult to learn accurate preference information. To alleviate the problem of insufficient effective recommendation signals caused by the sparsity of user-item interaction data, this paper proposes a recommendation algorithm framework named Multi-Representation Space Recommendation with Graph Contrastive Learning(MRSGCL). Compared with previous works that indirectly model higher-order relations through first-order relations, this paper additionally extracts a homogeneous graph, thereby enhancing the model's ability to represent preferences. To extract fine-grained preference information from user behavior, we propose an adaptive dual-representation space feature encoding module. This module learns the embeddings of users and items in heterogeneous spaces based on user-item (UI) interactions, and learns the embeddings of user-user (UU) and item-item (II) relationships in homogeneous spaces. Subsequently, to integrate the interest information learned from the two spaces, we propose an interest space alignment module. This module uses contrastive learning to bring the interest distributions of the two spaces closer together while preserving the distinct interest information learned from each space. Finally, we introduce an interest fusion module that combines the interest preferences from the two spaces using both adaptive and heuristic methods. We conducted extensive experiments on three public datasets, validating the effectiveness and robustness of our approach, achieving excellent experimental results. This demonstrates the scalability and high applicability of our method.
Tumor-associated macrophages (TAMs) are central orchestrators of immune evasion, therapeutic resistance, and clinical outcome across diverse solid tumors. However, the classical M1/M2 paradigm fails to capture the lineage diversity, spatial heterogeneity, and functional plasticity of TAMs within the tumor microenvironment. Unlike previous reviews that mainly summarize macrophage polarization or TAM subtype classification, this review focuses on how emerging single-cell and spatial multi-omics technologies redefine TAM ontogeny, phenotypic diversity, spatial organization, and therapeutic vulnerabilities. This review summarizes how fate-mapping, single-cell transcriptomics, and spatial multi-omics have contributed to the understanding of the dual origins of TAMs, including embryo-derived tissue-resident macrophages and monocyte-derived macrophages, and how these technologies further delineate transcriptionally and functionally distinct TAM subsets such as SPP1+, TREM2+, C1QC+, and MMP9+ TAMs with context-dependent roles in immune regulation, metastasis, and treatment resistance. We then integrate spatial transcriptomics, multiplex imaging, and radiomics evidence to outline TAM-enriched niches at hypoxic cores, invasive fronts, perivascular regions and tertiary lymphoid structures, emphasizing how these niches coordinate crosstalk among cancer cells, T cells, cancer-associated fibroblasts, endothelial cells and B cells. Emerging TAM-related gene signatures and multi-omics-based scores were further highlighted that predict response or resistance to radiotherapy, chemotherapy, anti-angiogenic therapy, and immune checkpoint blockade. Finally, we provide a forward-looking perspective on strategies to reprogram, deplete, or redirect specific TAM subsets, including CSF1/CSF1R and CCR2/CCL2 blockade, metabolic and epigenetic modulators, agonists of phagocytosis, and CAR-macrophage–based therapies. Overall, we propose that integrating single-cell and spatial omics with AI-assisted digital pathology and longitudinal sampling will enable TAM-informed patient stratification and rational design of combination immunotherapies, thereby accelerating the translation of TAM biology into precision oncology.
Arbitrary-oriented object detection remains a pivotal research focus due to its practical significance and inherent challenges. Existing methods often extend frameworks and sampling strategies designed for horizontal object detectors, which struggle to handle the arbitrary orientations, high aspect ratios, and diverse scales of oriented objects. To overcome these limitations, we propose a novel and efficient method for arbitrary-oriented object detection. This approach dynamically assigns prediction layers by object pixel area, then leverages wavelet transform-based energy weighting for bottom-up sample reassignment, optimizing feature representation for oriented targets. In addition, a robust framework integrates heatmap keypoint prediction on feature maps of a quarter-sized image, along with sparse predictions on other scales. By querying small-object regions within deep feature maps, a progressive top-down feature fusion strategy further enhances the perception of fine-grained details. Extensive evaluations on four benchmark datasets demonstrate the method's substantial improvements in detection performance, establishing its potential for broader applications in oriented object detection.
Peritoneal dialysis (PD) is an established home-based kidney replacement therapy but remains underutilized, partly due to limited patient knowledge. Meanwhile, TikTok is increasingly used for health information, but the quality of PD-related content on this platform remains uncertain. We conducted a cross-sectional evaluation of PD-related Douyin (the Chinese domestic counterpart of TikTok) videos retrieved on November 13, 2025. Video characteristics were recorded and videos were categorized by uploader type and content category. Two independent raters assessed each video using the Journal of the American Medical Association (JAMA) benchmark criteria, the modified DISCERN (mDISCERN), the Global Quality Score (GQS), the Patient Education Materials Assessment Tool for Audiovisual Materials (PEMAT-A/V), and Peritoneal Dialysis-Specific Scale (PDSS), including a Core PDSS subset. Among 145 included videos, most (143/145) received a JAMA score of 2. Median scores were 2.0 (1.0–3.0) for mDISCERN, 2.0 (2.0–3.0) for GQS, 80.0% (70.0%–95.4%) for PEMAT understandability, 66.7% (66.7%-100%) for PEMAT actionability, 5.0 (3.0–7.0) for PDSS and 2.0 (1.0–3.0) for Core PDSS. Scores differed significantly by uploader type and content category. Videos posted by patients and non-professional individuals had lower GQS, PDSS, and Core PDSS scores than those from professional individuals or hospital-affiliated organizations. Spearman correlation analysis revealed that likes were not significantly correlated with any quality score. Favorites and shares correlated positively with quality and reliability, whereas comments correlated negatively. Multiple linear regression analysis suggested that longer video duration was independently associated with higher quality and reliability as well as higher engagement (likes, favorites, and shares). Overall, PD-related content on Douyin is frequently insufficient in reliability and patient education value. Strengthening nephrologist-led production, platform governance, and actionable patient-centered education is warranted.
The translation of nucleic acid testing to point-of-care settings is hindered by the reliance on target amplification, which introduces complexity and contamination risks. Herein, we report a target-amplification-free assay for the direct detection of Staphylococcus aureus 16S rRNA, utilizing a cascaded DNAzyme and Nicking endonuclease reaction (DNECR). This system integrates a panel of target-specific multicomponent DNAzyme (MNAzyme) for primary recognition and signal amplification with a sterically blocked bipedal DNA walker (BDW) for cascade signal amplification. Upon target binding, the activated MNAzyme cleaves the blocker to initiate the BDW, which then traverses a spherical nucleic acid track via nicking endonuclease activity, generating amplified fluorescent signals. This cascaded design achieved a detection limit of 102 CFU/mL for cultured S. aureus with high specificity. Clinical validation using 12 patient sputum samples demonstrated 100% diagnostic sensitivity and specificity, confirming the potential of DNECR as a robust, amplification-free platform for rapid pathogen detection at the point of care.
Kinesin family genes (KIFs), a group of microtubule-associated motor proteins, have emerged as potential novel biomarkers in cervical cancer (CC). In the present study, a comprehensive bioinformatics analysis of KIF expression profiles was conducted using The Cancer Genome Atlas CC dataset and 23 KIFs with prognostic significance were identified. Non-negative matrix factorization based on their expression patterns revealed three distinct molecular subtypes of CC: C1, C2 and C3. To facilitate subtype prediction, a neural network model trained on KIF expression data was developed and validated in independent Mexican and Korean cohorts. Multi-omics characterization of the subtypes revealed distinct biological features: C1 was associated with downregulated oncogenic signaling; C2 exhibited activation of Hippo-YAP and VEGFR pathways; and C3 was characterized by Wnt signaling activation and an immune-silent phenotype. Predicted immunotherapy responses also varied across subtypes, with C1 patients anticipated to have the most favorable outcomes. Notably, KIF4A and KIF1A were identified as novel biomarker candidates specific to subtypes C2 and C3, respectively, and their expression patterns were validated in a Chinese CC cohort via immunohistochemistry, supporting their potential utility in prognostication and patient stratification. Overall, these findings provide new insights into the molecular heterogeneity of CC and highlight KIF genes as promising biomarkers for guiding personalized therapeutic strategies.
Multimodal recommendation systems have received widespread attention in recent years. The existing methods mainly treat multimodal features as a supplement to collaborative ID features to build alignment representation. However, these methods only use the user-item collaborative relation and itemitem latent relation to implicitly represent multimodal features without explicitly using multimodal content, which ignores the intrinsic semantic relations within different modal content and makes the model not pay insufficient attention to user preferences, causing not great recommendation performance. Considering this problem, we argue the complete fine-grained representation of multimodal content, which includes collaboration relation, latent relation, and semantic relation. To address this, this paper proposes the Joint Content Semantic relation learning with Mamba for Multimodal Recommendation (JCSMRec). First, we construct Semantic relation representation with Mamba ((SRM)-M-2) Module to explore the inter-modal correlation and provide an excellent aligned explicit representation for multimodal content. Then, we design a User preference-aware Module, according to user preference to guide learning the representation granularity of different modalities in the latent relations. Then, we conduct a joint representation to align and fusion multimodal features, which use the cyclic KL loss to align different semantic spaces and alleviate two joint losses to let user preference guide the learning process of the (SRM)-M-2 Module. This design effectively avoids over-reliance on ID features and enables a more comprehensive learning of textual and visual semantic information suitable for recommendation. Our method has achieved good results on multiple datasets and can be effectively inserted into other recommendation methods to improve results.
This study aimed to develop an effective predictive tool that combines radiomics and clinical information to predict the survival outcomes of patients with advanced non-small cell lung cancer (NSCLC) undergoing chemoimmunotherapy. Data were collected from 201 patients with advanced NSCLC who received first-line chemoimmunotherapy across three institutions: those from Centers I II (n = 164) were randomly split in a 7:3 ratio into training (n = 115) and validation (n = 49) cohorts, and those form Center III (n = 37) were designated as the external test cohort. The analysis was conducted using CT images and clinical data obtained before and after induction chemoimmunotherapy. We developed multiple intratumoral and peritumoral radiomics-based models, along with clinical prediction model that integrated patients’ baseline clinicopathological characteristics with plasma biomarker profiles, to predict progression-free survival (PFS). Based on expectations derived from prior established models, a stepwise backward elimination approach was utilized to select candidate submodels for the combined model construction. This combined model was internally validated using time-dependent ROC curves in training and validation sets and externally validated in the external test set. The combined model was constructed by integrating four candidate sub-models (DeltaSub, Clinical, P4mm, and Habitat) selected through the stepwise regression analysis. The combined model demonstrated superior performance compared to conventional models that utilized only clinical features, as well as Classical-Pre, Classical-Post, delta intratumor feature-based, and peritumor feature-based models. The combined model demonstrated satisfactory predictive performance across all three datasets, achieving a C-index of 0.849 (95
Many supervised learning-based facial expression recognition (FER) methods achieve good performance with the assistance of expression labels and a complex framework. However, there are inconsistent annotations in different expression datasets, making the above methods disadvantageous for new expression datasets or datasets with limited training data. The objective of this paper is to learn self-supervised facial expression features that enable the FER model not to rely on the annotation consistency of the different datasets. Most current self-supervised learning algorithms based on contrastive learning learn the representation by forcing different augmented views of the same image close in the embedding space, but they cannot cover all variances within a semantic class. We propose a heatmap neighbor contrastive learning (HNCL) method for FER. It treats the images corresponding to the heatmap nearest neighbors of expressions as other positives, providing more semantic variations than pre-defined augmented transformations. Therefore, our HNCL can learn better expression features covering more intra-class variances, improving the performance of the FER model based on self-supervised learning. After fine-tuning, HNCL with a simple framework achieves top-three performance on the in-the-lab datasets and even matches the performance of state-of-the-art supervised learning methods on the in-the-wild datasets.
Transformers-based trackers offer significant potential for integrating semantic interdependence between template and search features in tracking tasks. Transformers possess inherent capabilities for processing long sequences and extracting correlations within them. Several researchers have explored the feasibility of incorporating Transformers to model continuously changing search areas in tracking tasks. However, their approach has substantially increased the computational cost of an already resource-intensive Transformer. Additionally, existing Transformers-based trackers rely solely on mechanically employing multi-head attention to obtain representations in different subspaces, without any inherent bias. To address these challenges, we propose HEART (Historical Information Embedding And Subspace Re-weighting Tracker). Our method embeds historical information into the queries in a lightweight and Markovian manner to extract discriminative attention maps for robust tracking. Furthermore, we develop a multi-head attention distribution mechanism to retrieve the most promising subspace weights for tracking tasks. HEART has demonstrated its effectiveness on five datasets, including OTB-100, LaSOT, UAV123, TrackingNet, and GOT-10k.
Recently, diffusion model-based methods have utilized user interest features as guidance conditions to achieve stable generation results in sequential recommendation tasks. However, these models struggle to capture users' dynamic interests, as the interests of different users are often inconsistent. Moreover, the fixed number of interests predefined by existing models cannot adapt to the diverse preferences of users, making it difficult to further improve recommendation performance. To address these issues, we propose a novel generative sequential recommendation framework named ADIGRec (Adaptive User Dynamic Interest Guidance for Generative Sequential Recommendation), which adaptively focuses on users' dynamic interest features. Specifically, our framework combines users' dynamic features and inherent interest features encoded from historical sequences as new guidance conditions. Furthermore, we introduce a module that injects dynamic interest features into the noise item embeddings, enabling explicit interaction with the guidance conditions during the generation phase. This approach essentially fits the noise in the target space rather than the user preference space, leading to improved recommendation diversity. Additionally, we propose a novel regularization method to mitigate the impact of user interest routing collapse on the generation results. Extensive experiments on three publicly available datasets demonstrate that our method achieves superior performance compared to established baseline methods.
Most existing multimodal collaborative filtering recommendation (MCFRec) methods rely heavily on ID features and multimodal content to enhance recommendation performance. However, this paper reveals that ID features are effective but have limited benefits in multimodal collaborative filtering recommendation. Therefore, this paper systematically deconstruct the pros and cons of ID features: (i) they provide initial embedding but lack semantic richness, (ii) they provide a unique identifier for each user and item but hinder generalization to untrained data, and (iii) they assist in aligning and fusing multimodal features but may lead to representation shift. Based on these insights, this paper proposes IDFREE, an ID-free multimodal collaborative Filtering REcommEndation baseline. IDFREE replaces ID features with multimodal features and positional encodings to generate semantically meaningful ID-free embeddings. For ID-free multimodal collaborative filtering, it further proposes an adaptive similarity graph module to construct dynamic user-user and item-item graphs based on multimodal features. Then, an augmented user-item graph encoder is proposed to construct more effective user and item encoding. Finally, IDFREE achieves inter-multimodal alignment based on the contrastive learning and uses Softmax loss as recommendation loss. Basic experiments on three public datasets demonstrate that IDFREE outperforms existing ID-based MCFRec methods, achieving an average performance gain of 72.24
Pseudogenes are abundantly present in the human genome and are often thought of as nonfunctional nucleotide sequences, but a growing body of research suggests that pseudogenes can play important biological roles through a variety of pathways, and can be involved in the development of cancer. Lung cancer is one of the most prevalent cancers in the world and it is crucial to find new therapeutic strategies for the treatment of lung cancer. In recent years, studies on the effects of pseudogenes on lung carcinogenesis have increased rapidly. This has pointed to new directions in the diagnosis and treatment of lung cancer. Aim of this paper is to comprehensively discuss the role and influence of pseudogenes in the lung cancer, and the potential of pseudogenes as novel epigenetic targets in lung cancer diagnosis and prognosis and treatment, which is significant for realizing the clinical benefits of pseudogenes.
Whereas the role of OR7E47P in NSCLC is not clear, we explored the impact of OR7E47P on the prognosis of NSCLC patients and possible mechanisms through bioinformatics and experi-mental approaches. OR7E47P, underexpressed in NSCLC tumor tissues, is associated with better prognosis and enhanced immune benefits. Localization analyses showed OR7E47P may func-tions as a cytoplasmic processing pseudogene, regulating ROBO2 expression via the OR7E47P/miR-183-5p/ROBO2 pathway at the post-transcriptional level. Single-cell analyses re-vealed that ROBO2 reduces the tumor-supportive role of cancer-associated fibroblasts (CAFs) by promoting programmed cell death of FAP/TGFB1 + CAFs, enhancing immune infiltration, and suppressing the TGF-β/Smad signaling pathway in fibroblasts and NSCLC cells. Finally, we con-firmed our findings through experiments, where OR7E47P directly binds to miR-183-5p, with a binding site that allows for endogenous competition, leading to the upregulation of ROBO2 ex-pression and potentially regulate TGF-β secretion of fibroblast, thereby regulating the TME in NSCLC.