The Valence-Consistent Shift (VCS) effect–a foundational principle in framing theory–holds that consumers generally respond more favorably to positively framed information than to negatively framed information. While this effect is well established in human-delivered service contexts, its applicability to AI-powered service encounters remains underexplored. This study investigates whether, and under what conditions, the VCS effect works in AI-powered services. Five studies demonstrate that positive attribute framing generally works better than negative attribute framing in promoting AI-powered services. However, this effect reverses when the service is personalized (vs. standardized) and may attenuate when the decision-maker is the service provider (vs. service receiver). Drawing on regulatory focus theory and processing fluency, we further identify the underlying psychological mechanism driving these effects. This research extends the VCS effect to AI-powered service contexts by identifying boundary conditions under which the effect is strengthened, weakened, or reversed. It also provides evidence that processing fluency helps explain how attribute framing interacts with AI-powered service contexts to influence customer response.
Automatic Music Transcription (AMT) plays a critical role in converting music audio signals into editable and actionable symbolic scores, and holds significant practical applications in music domains such as music composition, digital music archiving, and automated music analysis. However, it still remain challenges to accurately identify notes, rhythms, and melodies, particularly in complex multi-track environments and dynamic note transcription tasks. Traditional AMT models typically rely on unimodal audio input, which struggles to effectively handle complex rhythmic structures and long-term dependencies between notes. To overcome these challenges, this paper proposes a novel AMT model based on a cross-modal Transformer architecture, named AMT-CMT. The model integrates audio features and structured score information by leveraging a dual-modality input mechanism and a cross-modal attention mechanism, to enable precise alignment between audio and score structures for high transcription performance especially in complex musical scenarios. The model incorporates a beat-depth-based note weighting mechanism that assigns hierarchical importance to notes according to their metrical positions, enhancing rhythmic structure modeling. It also employs a sparse Mixture-of-Experts (MoE) mechanism that dynamically routes inputs to specialized subnetworks, improving adaptability across diverse musical styles. Finally, we validate the superiority of our proposed AMT-CMT model by experiments on the MusicNet and URMP datasets with paired audio and MusicXML data. Experimental results indicate that the AMT-CMT model outperforms state-of-the-art models, with the improvement of 6.62% in Frame F1 and 6.63% in Onset F1. Furthermore, the model also improves the Onset+Offset F1 metric by 6.72% on the URMP and 6.67% on the MusicNet. These results fully demonstrates the robustness of our proposed model in complex musical scenarios for producing precise editable MusicXML files.
Diffusion models have demonstrated significant potential in medical image segmentation, yet they remain challenged in capturing complex fine-grained structures. Existing spatial-domain backbones exhibit bias towards low-frequency components, while the denoising process often leads to over-smoothing of high-frequency details, compromising boundary precision and cross-domain generalization. To overcome these limitations, we propose a Spatial-Frequency Aware Diffusion Network (SFA-DiffNet) that jointly models spatial and frequency representations, augmented with conditional guidance from foundation models. This design improves detail preservation and boundary accuracy in our experiments. Specifically, we introduce a Spatial-Frequency Aware Encoder (SFAE) that decomposes features into complementary spatial and frequency representations. It consists of a Spatial Aware Block (SAB) for contextual information learning, a Frequency Aware Block (FAB) for spectral cues exploration, and an Adaptive Fusion Block (AFB) for complementary fusion. Moreover, we develop a foundation model guidance mechanism that dynamically injects semantic priors via an Adaptive Feature Projector (AFP) and Semantic Alignment Fusion Block (SAFB), ensuring alignment between conditional and denoising features. Extensive experiments on multi-modal medical imaging datasets show that SFA-DiffNet achieves competitive performance against existing methods and shows potential for cross-domain generalization. The source code is available at GitHub.
Multimodal recommendation systems, leveraging multimodal data, demonstrate great potential in modeling user preferences and improving recommendation accuracy. However, due to the issues such as sparse interaction data, their performance typically deteriorates significantly, limiting their applicability. Existing works address this issue by using modality feature enhancement or self-supervised learning signal generation. The fusion features, which are derived from similar nodes with the original features, is suboptimal when enhancing user embedding features. To address these issues, we propose the Coupling Degree with Self-supervised Learning for Multimodal Recommendation (CDSSL). First, a novel concept is proposed named Coupling Degree (CD) in recommendation systems, which is utilized to enhance the representation of users and items by learning the coupling between nodes in the interaction graph. Furthermore, we guide the consistency learning of original interactions and generated interactions by combining modality information, thereby improving self-supervisory signals. The effectiveness of CDSSL was comprehensively assessed using four publicly available datasets, which is on average 6.12% higher than the current optimal baseline, demonstrating the effectiveness of the framework. The source code is available at https://github.com/Loyreem/CDSSL.
A new oligosaccharide ester, namely 3′-O-(Z)-3,4,5-trimethoxy-cinnamoyl-sucrose (1) along with nine known compounds (2–10), were isolated from the roots of Polygala fallax Hemsl. The structure of new compound was elucidated using 1D, 2D NMR, and HRESIMS methods. Furthermore, compounds 2–10 were isolated from the roots of Polygala fallax Hemsl for the first time. And the class of Lignan (8 and 10) were firstly reported in this species. Additionally, the antitumor activities of all compounds were evaluated.