Guangdong Institute of Science and Technology (simplified Chinese: 广东科学技术职业学院; traditional Chinese: 廣東科學技術職業學院院; pinyin: Guǎngdōng kēxué jìshù zhíyè xuéyuàn) is a provincial university located in the Tianhe District of Guangzhou City, Guangdong Province, China.
Voice messages have surged as an effective communication medium, offering convenience and rich paralinguistic cues. However, their reliance on audio playback often restricts message review in various situations. While transcriptions and teasers are helpful, they still require users to find a private place to listen to the audio. To address this limitation, we present VOVI, a voice-visualized messaging system that supports message review in environments where audio playback is impractical. Building on careful design rationale, we integrate visualization features into a smartwatch-based voice messaging interface. The system automatically detects speech nuances and proposes customizable visualized transcriptions. An user study with 20 participants showed that VOVI's transcriptions can capture speech content and paralinguistic cues, allowing senders to express nuances with less effort and helping receivers interpret them without audio. Our findings suggest that voice visualization has the potential to support voice message interactions and offer insights for designing future voice messaging systems.
A new anthraquinone-based tetra-benzimidazolium salt 1,8-bis2’-[2’’-(N-picoly-benzimidazoliumyl)ethyl]benzimidazoliumylethoxy-9,10-anthraquinone hexafluorophosphate (1) was prepared and characterized. Particularly, the recognition performance of H2PO4− using of compound 1 as a chemical sensor was investigated through fluorescence spectra, ultraviolet spectra, HRMS, 1H NMR titrations and IR spectra. The experimental results showed compound 1 has a good recognition ability for H2PO4−. One tetra-benzimidazolium salt 1 was prepared and characterized. The recognition of H2PO4− using 1 as a chemosensor was studied.
Robust grasping in cluttered environments remains an open challenge in robotics. While benchmark datasets have significantly advanced deep learning methods, they mainly focus on simplistic scenes with light occlusion and insufficient diversity, limiting their applicability to practical scenarios. We present GraspClutter6D, a large-scale real-world grasping dataset featuring: (1) 1,000 highly cluttered scenes with dense arrangements (14.1 objects/scene, 62.6% occlusion), (2) comprehensive coverage across 200 objects in 75 environment configurations (bins, shelves, and tables) captured using four RGB-D cameras from multiple viewpoints, and (3) rich annotations including 736K 6D object poses and 9.3B feasible robotic grasps for 52K RGB-D images. We benchmark state-of-the-art segmentation, object pose estimation, and grasp detection methods to provide key insights into challenges in cluttered environments. Additionally, we validate the dataset's effectiveness as a training resource, demonstrating that grasping networks trained on GraspClutter6D significantly outperform those trained on existing datasets in both simulation and real-world experiments. The dataset, toolkit, and annotation tools are publicly available on our project website: https://sites.google.com/view/graspclutter6d.
Recent progress in image-to-video (I2V) diffusion models has significantly advanced the field of generative inbetweening, which aims to generate semantically plausible frames between two keyframes. In particular, inference-time sampling strategies, which leverage the generative priors of large-scale pre-trained I2V models without additional training, have become increasingly popular. However, existing inference-time sampling, either fusing forward and backward paths in parallel or alternating them sequentially, often suffers from temporal discontinuities and undesirable visual artifacts due to the misalignment between the two generated paths. This is because each path follows the motion prior induced by its own conditioning frame. In this work, we propose Motion Prior Distillation (MPD), a simple yet effective inference-time distillation technique that suppresses bidirectional mismatch by distilling the motion residual of the forward path into the backward path. Our method can deliberately avoid denoising the end-conditioned path which causes the ambiguity of the path, and yield more temporally coherent inbetweening results with the forward motion prior. We not only perform quantitative evaluations on standard benchmarks, but also conduct extensive user studies to demonstrate the effectiveness of our approach in practical scenarios.
Diffusion model advances have enabled powerful text-guided image editing, but also raise ethical and legal risks such as deepfakes and unauthorized use. To prevent these risks, adversarial attack-based image immunization has emerged as a promising defense against AI-driven semantic manipulation. Yet, most existing approaches require image-specific optimization or additional neural networks at inference time, hindering scalability and practicality. In this paper, we propose the first universal adversarial perturbation-based image immunization framework that generates a single, image-agnostic adversarial perturbation specifically designed for diffusion-based editing pipelines. Inspired by UAP used in targeted attacks, our method aims to generate a UAP that induces diffusion models to misinterpret the input image as a specific semantic target. Simultaneously, it suppresses original content to misdirect the model's attention during editing, thereby effectively blocking unauthorized edits by overwriting the image's original semantics via the UAP. Extensive experiments show that our method, as the first universal immunization approach, significantly outperforms several baselines in the UAP setting. Notably, despite the inherent difficulty of universal perturbations, our method achieves competitive or superior performance compared to image-specific methods under a more restricted perturbation budget, while also exhibiting strong black-box transferability across diverse diffusion models.