Existing AI-based meeting summarization tools have enabled rapid generation of meeting notes, yet their reliability and user controllability remain limited. This paper explores human-AI collaboration for mobile meeting summarization and presents MeetSumAid, a multifunctional system that integrates summarization algorithms with an interactive user interface. The system is designed to support users in understanding, validating, and refining AI-generated summaries through natural interactions and flexible control mechanisms. By enabling real-time inspection, editing, and feedback, MeetSumAid facilitates reliable collaboration between humans and AI in dynamic meeting scenarios. A user study with 20 participants shows that MeetSumAid significantly improves summary quality, generation efficiency, and user-perceived reliability compared with baseline AI summarizers, while reducing cognitive load. Further analysis reveals how different interface components enhance users' engagement and confidence during collaboration. This work provides a practical step toward reliable and user-centered human-AI collaboration in mobile meeting summarization and offers actionable design implications for future intelligent collaborative systems.
Tibetan text recognition plays a key role in preserving the Tibetan language, religion, and traditions. While text recognition has made progress for high-resource languages, handwritten Tibetan character recognition remains difficult due to limited data and the lack of public large language models. Most existing datasets focus on printed or historical documents, as well as online handwriting data, but there are still few large offline handwritten Tibetan datasets. To solve this problem, we construct TibHCR, a large-scale offline handwritten character recognition dataset for the Tibetan language. To increase the diversity of the linguistic and font styles, more character categories and participants from 5 provinces in China are included. To collect and label the data efficiently, we introduce a grid sheet design, reducing manual annotation to just 1% of the samples. This design then allows for automatic data processing to extract each character sample and its corresponding label. The resulting TibHCR dataset contains 141,698 samples from 235 Tibetan writers, covering 47 character classes. We evaluate TibHCR using two recognition models: a convolutional recurrent neural network (CRNN) and a cross-lingual fine-tuning method, on a Chinese pretrained model using the PP-OCRv4 architecture to adapt Tibetan data. The results show that both models can recognize handwritten Tibetan characters efficiently, with an accuracy of 99.48% for CRNN and 99.70% for the fine-tuning method. The TibHCR dataset is publicly available at https://huggingface.co/datasets/qixiaoke/TibHCR.
Traditional contract review becomes increasingly time-consuming when attorneys face a surge in the number of contracts. Artificial Intelligence (AI) offers a solution to enhance the efficiency of reviews, but its uncertainty raises concerns among legal professionals. In this paper, we explore innovative clues and interaction design for trust calibration through a scenario-centric approach. Conducting formative experiments and semi-structured interviews, we conducted a contextual investigation with 24 attorneys, uncovering mismatches and trust calibration challenges between commercial AI tools and manual review processes in practical use. Based on these findings, we collaboratively designed seven key design components with attorneys and developed the ContractMind system prototype using a Wizard-of-Oz design approach. Through evaluation with 16 attorneys, we calculated and analyzed the differences between participants' perceived trustworthiness and AI system capabilities. Compared to commercial artificial intelligence tools, our system is more conducive to trust calibration, enabling them to smoothly review contracts and make informed decisions. This work takes the first step in exploring trust calibration design for AI contract review tools and designs and evaluates ContractMind system prototype. We also provide design considerations for future AI contract review tools and trust calibration.
Rate adaptation in LoRa communications is crucial for improving the channel throughput by adjusting the data rate according to varying channel conditions. Existing methods typically operate at the packet or symbol level, which limits their ability to achieve fine-grained rate adaptation. In this paper, we propose ILoRa, an Interleaving-driven partial transmission method that automatically adjusts transmission rates according to real-time channel conditions. To be specific, we first introduce intra-symbol interleaving that leverages a progressive inorder traversal method to determine the transmission order within a symbol. Then inter-symbol interleaving is applied to coordinate the order across symbols. To manage the interleaving-induced partial transmission and improve communication performance under noisy conditions, we employ a multitask convolutional recurrent neural network (MT-CRNN). This network leverages advanced data augmentation methods to further enhance channel robustness: time-spectral augmentation to mitigate information loss and synthetic noisy data to simulate various channel conditions. Extensive experimental results demonstrate that ILoRa significantly enhance transmission efficiency while maintaining reliable performance even in challenging environments.
AI-driven advancements in speech synthesis and voice conversion, now are able to convincingly emulate human speech, have made a growing challenge for investigators and the judicial system to discern between genuine and artificially generated audio. Creating an effective audio deepfake detector necessitates a large-scale and high-quality data. Existing datasets focus on monolingual data for high-resource languages. In this paper, we have constructed a multi-lingual multi-speaker audio deepfake dataset, named MADD. The sources of MADD are derived from Common Voice and Gigaspeech2. Leveraging various deep speech synthesis and voice conversion technologies across 6 languages, the MADD dataset comprises a collection of 129,990 deep synthesis utterances, with a total duration of 155.66 hours from 288 speakers.
Low-resource text plagiarism detection faces a significant challenge due to the limited availability of labeled data for training. This task requires the development of sophisticated algorithms capable of identifying similarities and differences in texts, particularly in the realm of semantic rewriting and translation-based plagiarism detection. In this paper, we present an enhanced attentive Siamese Long Short-Term Memory (LSTM) network designed for Tibetan-Chinese plagiarism detection. Our approach begins with the introduction of translation-based data augmentation, aimed at expanding the bilingual training dataset. Subsequently, we propose a pre-detection method leveraging abstract document vectors to enhance detection efficiency. Finally, we introduce an improved attentive Siamese LSTM network tailored for Tibetan-Chinese plagiarism detection. We conduct comprehensive experiments to showcase the effectiveness of our proposed plagiarism detection framework.
Alzheimer's disease (AD) is considered as one of the leading causes of death among people over the age of 70 that is characterized by memory degradation and language impairment. Due to language dysfunction observed in individuals with AD patients, the speech-based methods offer non-invasive, convenient, and cost-effective solutions for the automatic detection of AD. This paper systematically reviews the technologies to detect the onset of AD from spontaneous speech, including data collection, feature extraction and classification. First the paper formulates the task of automatic detection of AD and describes the process of data collection. Then, feature extractors from speech data and transcripts are reviewed, which mainly contains acoustic features from speech and linguistic features from text. Especially, general handcrafted features and deep embedding features are organized from different modalities. Additionally, this paper summarizes optimization strategies for AD detection systems. Finally, the paper addresses challenges related to data size, model explainability, reliability and multimodality fusion, and discusses potential research directions based on these challenges.
Proceeding from the actual problems of Chinese minority language information processing, this paper establishes a Tibetan-Chinese cross-language text plagiarism detection corpus containing 150,000 sentence pairs, based on SemEval 2014 English evaluation corpus and data enhancement method, to solve the problem of lack of corpus in Tibetan-Chinese cross-language text plagiarism detection. This dataset provides fundamental basis for Tibetan-Chinese cross-language text plagiarism detection. Also the dataset can be used in Tibetan-Chinese semantic computing and other natural language processing tasks. In addition, data enhancement method in the process of data set construction also provides a solution for other low-resource languages to solve the problem of lack of corpus in natural language processing tasks.
Supporting a massive amount of Internet of Things applications requires a large pool of spectrum. DSM is a promising ecosystem to improve the spectrum efficiency. In the era of LoRaWAN, the physical hardware constraints, along with the bandwidth-hungry applications pose new challenges. In this article, we investigate a novel deep-reinforcement-learning-based spectrum-sharing paradigm, termed Intelligent Overlapping, that explores partially overlapping channels for concurrent spectrum access in LoRaWAN. Our key insight is to leverage the coding redundancy to expand the available spectrum without complicated data processing algorithms. In particular, we learn the extra coding redundancy from the data on the non-overlapping spectrum via a deep-Q-learning network, and we apply such redundancy to recover the data on the overlapping spectrum. In the Media Access Control layer, we predict the channel condition and strategically learn and assign the appropriate overlapping portion to the concurrent access end devices. In the Physical layer, we harness interleaving to randomize the mutual interference to ensure that all the data remains decodable. Simulation results demonstrate that Intelligent Overlapping greatly improves the spectrum efficiency with a fast convergence rate compared to the conventional DSM mechanisms.
针对机器人的自动对准问题,提出一种基于点线特征的解耦视觉伺服控制方法.所提方法以点和直线作为图像特征,并利用图像特征的交互矩阵解耦姿态控制和位置控制,实现六自由度对准.首先利用直线及其交互矩阵设计姿态控制律,以消除旋转偏差;然后利用点及其交互矩阵设计位置控制律,以消除位置偏差;最后实现机器人末端目标的自动对准.在对准控制过程中,基于执行的相机运动量以及相机运动前后特征的变化量,可实现对深度的在线估计.另外,还设计了监督器对相机的运动速度进行调节,从而确保特征一直处于相机视野当中.在Eye-in-Hand机器人平台上,分别用所提方法和传统的基于图像的视觉伺服方法实现了机器人的六自由度对准.所提方法经过16步实现了机器人的自动对准,对准结束时机器人末端位姿的最大平移误差为3.26 mm,最大旋转误差为0.72°.相较于对比方法,该方法的控制过程更加高效,控制误差收敛更快,对准误差更小.实验结果表明,所提方法可以实现快速高精度的自动对准,能够提高机器人操作的自主性和智能化水平,有望应用于目标跟踪、拾取和定位、自动化装配、焊接、服务机器人等领域.
个性化的头相关传输函数(head-related transfer function,HRTF)可以有效改善空间音频质量.针对个性化HRTF难以精确获得的问题,提出了一种基于层级集成的个性化空间音频生成方法.该方法通过三个模型逐层建立个性化HRTF中的定位信息.首先,采用高斯混合模型建立用户无关的共用模型.然后,采用自编码器获得与用户有关的HRTF的隐表示,利用深度神经网络在人体生理参数与HRTF的隐表示之间建立非线性映射,得到用户有关的个性化模型.为了尽可能恢复个性化HRTF细节信息,对上述模型降维过程中的残差进行线性建模,得到残差模型.对于目标用户,任意空间位置处的个性化的HRTF可以通过集成三个层次下的模型获得,用于生成三维空间音频.最终,实验结果表明,提出的算法可以有效降低HRTF频谱损失,提升对个性化HRTF的预测性能.
Due to the lack of public datasets, few researches focus on speech translation in minority languages. To this end, this paper constructs a dataset of Mongolian-Chinese speech translation, named as NMLR-Mon2Chs ST. The dataset consists of Mongolian speech, Mongolian and Chinese text. First, Mongolian speech were obtained from 36 Mongols aged between 20 and 25 by recording on their mobile phones. Then, the corresponding Chinese texts were annotated by professionals. In order to make sure the quality of the dataset, the preprocessing was done, such as removing the quiet speech, resampling, and normalization. As a result, a total of 25 hours of high-quality data are obtained, and the average duration of audio in the dataset is 4.2 seconds. The establishment of this dataset allows researchers access to speech translation for minority languages.
Spatial audio has attracted more and more attention in the fields of virtual reality (VR), blind navigation and so on. The individualized head-related transfer functions (HRTFs) play an important role in generating spatial audio with accurate localization perception. Existing methods only focus on one database, and do not fully utilize the information from multiple databases. In light of this, a pre-trained-based individualization model is proposed to predict HRTFs for any target user in this paper, and a real-time spatial audio rendering system built on a wearable device is implemented to produce an immersive virtual auditory display. The proposed method first builds a pre-trained model based on multiple databases using a DNN-based model combined with an autoencoder-based dimensional reduction method. This model can capture the nonlinear relationship between user-independent HRTFs and position-dependent features. Then, fine tuning is done using a transfer learning technique at a limit number of layers based on the pre-trained model. The key idea behind fine tuning is to transfer the pre-trained user-independent model to the user-dependent one based on anthropometric features. Finally, real-time issues are discussed to guarantee a fluent auditory experience during dynamic scene update, including fine-grained head-related impulse response (HRIR) acquisition, efficient spatial audio reproduction, and parallel synthesis and playback. These techniques ensure that the system is implemented with little computational cost, thus minimizing processing delay. The experimental results show that the proposed model outperforms other methods in terms of subjective and objective metrics. Additionally, our rendering system runs on HTC Vive, with almost unnoticeable delay.
Recent years have witnessed a steep grow in the multimedia-oriented Internet of Things (IoT) over vehicular networks. Huge volume of multimedia traffic generated from the in-built IoT devices should be delivered among vehicles and immediate surroundings in real time. However, as network nodes with higher mobility, vehicles often experience more unpredictable wireless channels. Such time-frequency diversity poses substantial challenges to achieve pervasive and real-time multimedia connectivity. The hurdle lies in the inability of automatically approaching the subcarrier level channel variations in the existing video codecs. With coarse-grained traffic delivery rate, video decoding fails, and intermittent connection occurs. To break this stalemate, we propose a fine-grained wireless video streaming strategy, namely iCast, that intelligently achieves the most appropriate data rate and frame protection for multimedia traffic in highly mobile vehicular environments. The insight of iCast is a simple joint source-channel rateless code. It reaps the benefits of the frequency diversity to provide fine-grained data rate for the channel in conjunction with suitable protection for the source. Our experiments show that, by harnessing frequency diversity in mobile environments, iCast outperforms the existing competitive wireless video delivery schemes by up to 5 dB peak signal-to-noise ratio.
Individualized head-related transfer functions (HRTFs) play an important role in accurate localization perception. However, it is a great challenge to efficiently measure continuous HRTFs for each person in full space. In this paper, we propose a parameter-transfer learning method termed PTL to obtain individualized HRTFs based on a small set of HRTF measurements. The key idea behind PTL is to transfer a HRTF generation model from other database to a target individual. To this end, PTL first pretrains a deep neural network (DNN)-based universal model on a large database of HRTFs with the assist of domain knowledge. Domain knowledge is used to generate the input features derived from the solution to sound wave propagation equation at the physical level, and to design the loss function based on the knowledge of objective evaluation criterion. Then, the universal model is transferred to a target individual by adapting the parameters of a hidden layer of DNN with a small set of HRTF measurements. The adaptation layer is determined by experimental verification. We also conduct the objective and subjective experiments, and the results show that the proposed method outperforms the state-of-the-arts methods in terms of LSD and localization accuracy.
Spherical harmonic (SH)-based methods have been proposed for modeling head-related transfer functions (HRTFs) and yielded an encouraging performance level in terms of log-spectral distortion (LSD). However, most of these techniques model HRTFs on a sphere, and rarely exploit the correlation relationship of HRTFs from different distances, and as a consequence HRTF extrapolation on unmeasured distances becomes a great challenge. Motivated by this, this paper proposes a distance-dependent SH-based model termed DSHM for HRTF representation. DSHM extends the SH-based model by adding a radial part of spherical Fourier-Bessel transform (SFBT). By utilizing a radial correlation between distances, the proposed model has capable of efficient representation for HRTFs over the whole space. As a result, it is feasible to interpolate or extrapolate an HRTF on an unmeasured position. The experimental results show that DSHM achieves a lower LSD when comparing with the conventional SH-based method.
Head-related transfer functions (HRTFs) describe the propagation of sound waves from the sound source to ear drums, which contain most of information for localization. However, HRTFs are highly individual-dependent, and thus because of the difference of anthropometric features between subjects, individualization of HRTFs is a great challenge for accurate localization perception in virtual auditory displays (VAD). In this paper, we propose a sparsity-constrained weight mapping method termed SWM to obtain individual HRTFs. The key idea behind SWM is to obtain optimal weights to combine HRTFs from the training subjects based on the relationship of anthropometric features between the target subject and the training subjects. To this end, SWM teams two sparse representations between the target subject and the training subjects in terms of anthropometric features and HRTFs, respectively. A non-negative sparse model is used for this purpose when considering the non-negative property of the anthropometric features. Then, we build a mapping between the two weight vectors using a nonlinear regression. Furthermore, an iterative data extension method is proposed in order to increase training samples for mapping model. The objective and subjective experimental results show that the proposed method outperforms other methods in terms of log-spectral distortion (LSD) and localization accuracy.
Many methods have been proposed for modeling head-related transfer functions (HRTFs) and yield a good performance level in terms of log-spectral distortion (LSD). However, most of them utilize linear weighting to reconstruct or interpolate HRTFs, but not consider the inherent nonlinearity relationship between the basis function and HRTFs. Motivated by this, a domain knowledge-assisted nonlinear modeling method is proposed based on bottleneck features. Domain knowledge is used in two aspects. One is to generate the input features derived from the solution to sound wave propagation equation at the physical level, and the other is to design the loss function for model training based on the knowledge of objective evaluation criterion, i.e., LSD. Furthermore, with utilizing the strong representation ability of the bottleneck features, the nonlinear model has the potential to achieve a more accurate mapping. The objective and subjective experimental results show that the proposed method gains less LSD when compared with linear model, and the interpolated HRTFs can generate a similar perception to those of the database.
Many methods have been proposed for modeling head-related transfer functions (HRTFs) and yield a good performance level in terms of log-spectral distortion (LSD). However, most of them utilize linear weighting to reconstruct or interpolate HRTFs, but not consider the inherent nonlinearity relationship between the basis function and HRTFs. Motivated by this, a domain knowledge-assisted nonlinear modeling method is proposed based on bottleneck features. Domain knowledge is used in two aspects. One is to generate the input features derived from the solution to sound wave propagation equation at the physical level, and the other is to design the loss function for model training based on the knowledge of objective evaluation criterion, i.e., LSD. Furthermore, with utilizing the strong representation ability of the bottleneck features, the nonlinear model has the potential to achieve a more accurate mapping. The objective and subjective experimental results show that the proposed method gains less LSD when compared with linear model, and the interpolated HRTFs can generate a similar perception to those of the database.
Several functional models for head-related transfer function (HRTF) have been proposed based on spherical harmonic (SH) orthogonal functions, which yield an encouraging performance level in terms of log-spectral distortion (LSD). However, since the properties of subbands are quite different and highly subject-dependent, the degree of SH expansion should be adapted to the subband and the subject, which is quite challenging. In this paper, a sparse spherical harmonic-based model termed SSHM is proposed in order to achieve an intelligent frequency truncation. Different from SH-based model (SHM) which assigns the degree for each subband, SSHM constrains the number of SH coefficients by using an l_1 penalty, and automatically preserves the significant coefficients in each subband. As a result, SSHM requires less coefficients at the same SD level than other truncation methods to reconstruct HRTFs. Furthermore, when used for interpolation, SSHM gives a better fitting precision since it naturally reduces the influence of the fluctuation caused by the movement of the subject and the processing error. The experiments show that even using about 40% less coefficients, SSHM has a slightly lower LSD than SHM. Therefore, SSHM can achieve a better tradeoff between efficiency and accuracy.