The two main challenges in period dating Chinese historical text are the enormous size of the text and its intrinsic monosyllabic nature. In this study, we propose an efficient framework to enhance the Chinese historical text period classification task. This paper introduces a framework, TGGGS (TopicGraph Guwen GraphSaint), by combining four different model architectures. Using a combination of Topic, Graph, Neural Networks (NNs), and Transformer architectures together with a Subgraph-wise sampling technique for optimal feature extraction, we can effectively improve the overall performance of a Graph Neural Network (GNN). The overall model framework consists of a Topic and Graph architecture (TopicGraph), followed by a Subgraph-wise sampling technique (GraphSAINT), and a multiclass Transformer-based classifier (GuwenBERT and NNs). The inclusion of the Subgraph-wise sampling technique increases the performance in the model without reducing the efficiency of the model. Moreover, the original Chinese historical text is inputted without any preprocessing techniques where both punctuation marks and stop words are kept intact. In this way, we preserved the original contextual information of a Chinese historical text. Experimental results demonstrate the effectiveness of our proposed approach.
This study presents a deep learning framework optimizing 3D clothing models for VR, using a CNN to significantly reduce the triangle count of models from DeepFashion3D and CAP-UDF datasets. Achieving a balance between efficiency and detail, it cuts triangle count from over 160,000 to below 4,000, maintaining high DPI. The approach automates optimization, promising scalability and efficiency in VR fashion, setting a foundation for future 3D content development, enhancing virtual garment realism and interactivity.
In recent years, there has been a rapid growth in the volume of textual data generated from various sources, including industries, news media, and social media, across various fields worldwide. It contains valuable information and knowledge, but its sheer volume requires effective summarization techniques to make it useful. Text summarization has thus become an important technology for distilling large amounts of information and capturing key insights. With advancements in large language models (LLMs), most existing abstractive summarization models leverage pretrained models such as BART. These models are based on Transformer architectures and undergo unsupervised pretraining, followed by fine-tuning on specific datasets for downstream tasks, particularly text summarization. These models have demonstrated superiority in this field. Real-world data often contains noise, and most pretrained language models suffer from exposure bias and a discrepancy between their training objectives and real evaluation metrics. In this paper, we introduce sentence-level data augmentation to simulate real-world text variations and enhance the model's denoising capability. Furthermore, contrastive learning is employed to learn to distinguish the differences between the original article with candidate summaries and among candidate summaries themselves. Finally, the highest-scoring article is output as the resulting summary. The experimental results on the CNN/DailyMail and XSum datasets indicate that this model shows improvement in generating text summaries compared to pretrained models and existing methods with contrastive learning. When comparing to the multitask learning framework BRIO variant with contrastive loss, our proposed method achieves a better ROUGE-1 of 48.52 and comparable ROUGE-2/L of 24.84 and 39.78. This shows the potential of our proposed method. Further investigation is needed to verify the performance in real-world texts in various domains.
This article examines the incorporation of the Shopping Assistance Automatic Suggestion (SAAS) model into Virtual Reality (VR) environments in order to improve the online shopping experience. The SAAS model employs sophisticated deep learning methods to offer customized product recommendations, which are conveyed by non-player characters (NPCs) via voice-based interactions. Our goal is to develop an interactive shopping experience that replicates real-life interactions by integrating AI-powered recommendations with immersive VR technology. We gather and standardize data from several open commerce databases, such as Amazon Product and Customer Reviews. The SAAS model, in conjunction with GPT-3, BERT, and T5, undergoes training and testing to evaluate its effectiveness across multiple criteria. The results demonstrate that the SAAS model surpasses other models in delivering contextually aware and pertinent recommendations. The integration process outlines the specific steps involved in capturing, processing, and transforming user interactions in virtual reality (VR) into vocal suggestions provided by non-player characters (NPCs). This strategy improves customization and utilizes the immersive features of virtual reality to effectively engage people. The results of our research establish a higher standard for e-commerce, with the goal of enhancing the user experience of online purchasing by making it more instinctive, engaging, and pleasurable.
Information sharing on social media has become a common practice for people around the world. Since it is difficult to check user-generated content on social media, huge amounts of rumors and misinformation are being spread with authentic information. On the one hand, most of the social platforms identify rumors through manual fact-checking, which is very inefficient. On the other hand, with an emerging form of misinformation that contains inconsistent image–text pairs, it would be beneficial if we could compare the meaning of multimodal content within the same post for detecting image–text inconsistency. In this paper, we propose a novel approach to misinformation detection by multimodal feature fusion with transformers and credibility assessment with self-attention-based Bi-RNN networks. Firstly, captions are derived from images using an image captioning module to obtain their semantic descriptions. These are compared with surrounding text by fine-tuning transformers for consistency check in semantics. Then, to further aggregate sentiment features into text representation, we fine-tune a separate transformer for text sentiment classification, where the output is concatenated to augment text embeddings. Finally, Multi-Cell Bi-GRUs with self-attention are used to train the credibility assessment model for misinformation detection. From the experimental results on tweets, the best performance with an accuracy of 0.904 and an F1-score of 0.921 can be obtained when applying feature fusion of augmented embeddings with sentiment classification results. This shows the potential of the innovative way of applying transformers in our proposed approach to misinformation detection. Further investigation is needed to validate the performance on various types of multimodal discrepancies.
Video summarization aims to select a subset of video segments that best capture the video storyline. Our study seeks to train an encoder to transform the raw frame features extracted from pre-trained CNN models into representations that embody importance and guide the selection of the video segment. Our main idea is to use graph modeling and attention mechanisms to train the encoder adversarially. The graph representation enables the model to learn the relationship among frames, revealing the intrinsic structure of a video. The attention mechanism allows the model to capture the magnitude of these relationships. In the proposed model, an attention-based encoder is trained using a graph-based generator that reconstructs videos using the encoded features and a discriminator that guides the generator, distinguishing the original and reconstructed video. Thus, by leveraging graph attention and refining mechanisms, the proposed model offers distinct advantages over existing methods, including enhanced summarization accuracy, improved preservation of temporal coherence, and the ability to capture complex semantic linkages within video content. These advancements are substantiated through a comprehensive ablation study, which demonstrates the efficacy of our model using various evaluation metrics - Kendall and Spearman coefficients. The proposed model is evaluated on TVSum and SumMe datasets and achieves results on par with supervised models that used similar encoders and achieved state-of-the-art results compared to other unsupervised models.
Objectives: The pulse wave fluctuations during pregnancy exhibit distinct characteristics that represent both the pregnancy and the fetus. However, these fluctuations, which contain more frequency domain information, have not yet been studied for diagnostic purposes. Our objective was to quantify these pulse waveforms and gather empirical evidence regarding their relationship with both pregnancy and fetal sex. Methods: We collected wrist pulse data from 36 pregnant women (420 datasets) and 50 nonpregnant women using a pressure sensor. We applied Fourier transformation to analyze these pulse waveforms and extract their harmonics in frequency domain. Subsequently, we conducted a statistical analysis to compare these harmonics between pregnant and nonpregnant women to estimate the differences. To validate our findings, we employed machine-learning techniques to assess the utility of these harmonics. Results: We observed a significant increase in the first and second harmonics in the right hand of the pregnant women, with a distinct pattern observed for these harmonics in the left hand. Furthermore, the differences in the top three harmonics between both hands were more evident, particularly in male fetuses, as we analyzed the monthly changes in harmonic differences, revealing stronger magnitudes in the early months. Additionally, our model accurately predicted fetal sex after 12 weeks over 70% accuracy using decision tree techniques. Conclusions: The pulse waveforms of the pregnant women displayed distinct signals, and the observed differences in harmonics between both hands were notable, especially with regard to fetal sex. This evidence could inform prenatal care practices and help research the maternal-fetal relationship.
With the development of mobile Web technologies, people can easily seek advice from social media before making purchases or decisions. Some companies employ expert writers to fabricate reviews or use automated techniques to improve the appeal of their products or services, or to undermine the credibility of their rivals. This obstructs the detection of fake reviews and reviewers. This paper proposes a novel graph neural network-based framework for detecting spammers, who originate fake reviews in discussion forums to capture information from different social network combinations in various subgraphs. These subgraphs include a complete social context graph, homogeneous user–user subgraph, and heterogeneous user–post subgraph. A novel two-stage architecture with focal loss was designed to create a training model. This model can be applied to solve the issue of imbalance data classification. The proposed framework was applied to evaluate a ground truth dataset collected from an actual fraudulent review event on a discussion forum. The experimental results show that this aggregate social context representation method can be effectively applied to detect fake reviewers.
Virtualization technologies are still growing bigger and faster. Despite the greatness of its advancement, the costume industry is still very accessible when it comes to real trials. Off-the-shelf stuff are inadequate details for the desired individual to assess its in-depth utility for each garment trying on for a second, including custom stuff are much harder to try out right away. To this end, 2D image-based 3D reconstruction inclusive of touchable-virtualized space is accessible easier to stuff details for mans' decision making in purchasing. We establish the overall end-to-end pipeline from reconstruction until visualization for one instance to be triable on its stuff for a moment. As an expectation, our proposed approach can bring objects into the experimental area and use them immediately without any obstacle.
Digital twin technologies are still developing and are being increasingly leveraged to facilitate daily life activities. This study presents a novel approach for leveraging the capability of mobile devices for photo collection, cloud processing, and deep learning-based 3D generation, with seamless display in virtual reality (VR) wearables. The purpose of our study is to provide a system that makes use of cloud computing resources to offload the resource-intensive activities of 3D reconstruction and deep-learning-based scene interpretation. We establish an end-to-end pipeline from 2D to 3D reconstruction, which automatically builds accurate 3D models from collected photographs using sophisticated deep-learning techniques. These models are then converted to a VR-compatible format, allowing for immersive and interactive experiences on wearable devices. Our findings attest to the completion of 3D entities regenerated by the CAP–UDF model using ShapeNetCars and Deep Fashion 3D datasets with a discrepancy in L2 Chamfer distance of only 0.089 and 0.129, respectively. Furthermore, the demonstration of the end-to-end process from 2D capture to 3D visualization on VR occurs continuously.
Online reviews affect consumers' buying decisions. When astroturfing happens, the posted fake reviews not only cause confusion to the consumers, but in extreme cases, harm the reputation of others. Identifying the original posters of fake review in a discussion forum from the interactions with the peers in a collusive effort is essential to an overall detection strategy. However, existing detection approaches that focus on the mechanic characteristics of the posted reviews may fall short, since fake reviews cause more harm when astroturfing campaign happens involving multiple accomplices. We propose a detection framework with graph neural network, which incorporates the original perpetrator's stylometric patterns and relationships with other accomplices. The framework was tested against the data collected from real incidents. Multiple deep learning models with fusion techniques were tested. Managerial and theoretical implications are provided.
Background and aim: Acupuncture has been criticized as a theatrical placebo for the sham effect. Unfortunately, sham tests used in control groups in acupuncture studies have always ignored the underlying biophysical factors, including resonance involved in acupuncture points and meridians.Experimental procedure: In this study, the effects of sham acupuncture at Tsu San Li (St-36) were examined by analyzing noninvasive 30-sec. recordings of the radial arterial pulses for 3 groups of patients treated with different probes (blunt, sharp, and patch) on the superficial skin of the acupuncture point. The 3 groups were then treated with the sharp probe for 3 different periods (16, 30, and 50 s). Then we compared the harmonics of the radial arterial pulse after Fourier transformation before and after the treatment. Results: Our results indicated that different probes have effects similar to needle insertion at Tsu San Li. Meanwhile, the harmonic effect of the sharp probe strengthened as time increased. Conclusions: This study revealed that the meridian effect of sham testing from mechanical stimulation, even from simple touch, on an acupuncture point, should not be overlooked. Thus, even simple touch can be added to electrical or laser acupuncture.(c) 2023 Center for Food and Biomolecules, National Taiwan University. Production and hosting by Elsevier Taiwan LLC. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/ licenses/by-nc-nd/4.0/).
Solving math word problems is a popular topic in natural language processing. We not only need to classify the grammatical structures in the questions, but also understand the mathematical logic expressed between words. Errors in semantic understanding may lead to the failure to generate correct solution equations. Thus, the correct answer cannot be calculated. Previous studies mainly used sequence-to-sequence recurrent neural networks (RNNs) to obtain meaning in words, or combined graph neural networks to capture more information in questions to achieve better results. In addition, recent studies also showed that, a tree-based decoder leads to better results than a decoder of RNNs. In this paper, we propose to combine transformers and tree-based decoders for solving math word problems. Firstly, we use a transformer encoder to read math word problems, whose outputs are given to two different decoders, including Transformer decoder, and a treebased decoder. Secondly, from the answer equations generated from the two decoders, the better solution is selected. The experimental results on the two commonly used math word problem datasets, MAWPS and ASDiv-A, show that our model achieves 89.3% and 81.9% accuracy, which are 2.2% and 4.2% higher than the vanilla transformer model, respectively. For the MAWPS dataset the performance is comparable to state-of-the-art model Graph2Tree. This shows the effectiveness of our proposed method.
Sharing information on social media has become a part of people's daily lives. However, without centralized management on user generated contents, increasing amount of rumors and misinformation are being spread in social media. On the one hand, most of the social platforms debunk rumors with manual verification by fact checking organizations, which are very inefficient. On the other hand, since it's common for rumors to contain inconsistent images and texts, it would be useful if we could compare the semantics between multimodal contents in the same post for rumor detection. In this paper, we propose to check multimodal content consistency with transformers and self-attention-based Bi-GRU networks for rumor detection. Firstly, image semantic contents are extracted by image captioning module to generate captions. Then, the generated captions are semantically compared with texts using transformers for veracity assessment. Finally, Multi-cell bi-directional Recurrent Neural Networks (Bi-RNNs) with self-attention mechanism are used to find word dependency and learn the most important features for rumor detection. From the experimental results on tweets, the best F1-score of 0.92 can be obtained for our proposed approach to multimodal veracity assessment. This shows the potential of our proposed method in rumor detection. Further investigation is needed to verify the performance using different multimodal features.
Nowadays short texts can be widely found in various social data in relation to the 5G-enabled Internet of Things (IoT). Short text classification is a challenging task due to its sparsity and the lack of context. Previous studies mainly tackle these problems by enhancing the semantic information or the statistical information individually. However, the improvement achieved by a single type of information is limited, while fusing various information may help to improve the classification accuracy more effectively. To fuse various information for short text classification, this article proposes a feature fusion method that integrates the statistical feature and the comprehensive semantic feature together by using the weighting mechanism and deep learning models. In the proposed method, we apply Bidirectional Encoder Representations from Transformers (BERT) to generate word vectors on the sentence level automatically, and then obtain the statistical feature, the local semantic feature and the overall semantic feature using Term Frequency-Inverse Document Frequency (TF-IDF) weighting approach, Convolutional Neural Network (CNN) and Bidirectional Gate Recurrent Unit (BiGRU). Then, the fusion feature is accordingly obtained for classification. Experiments are conducted on five popular short text classification datasets and a 5G-enabled IoT social dataset and the results show that our proposed method effectively improves the classification performance.
User-generated contents in social media are not verified before being posted. They could bring many problems if they were misused. Among various types of rumors, the authors focus on the type in which there's mismatch between images and their surrounding texts. They can be detected by multimodal feature fusion in RNNs with attention mechanism, but the relations between images and texts are not well-addressed. In this paper, the authors propose to improve rumor detection by image captioning and RNNs with self-attention. Firstly, they utilize the idea of image captioning to translate images into the corresponding text descriptions. Secondly, these caption words are represented by word embedding models and aggregated with surrounding texts using early fusion. Finally, multi-cell bi-directional RNNs with self-attention are used to learn important features to identify rumors. From the experimental results, the best F-measure of 0.882 can be obtained, which shows the potential of our proposed approach to rumor detection. Further investigation is needed for data in larger scale.
Stock markets are often influenced by various factors which makes it very challenging to predict. Machine learning and deep learning models are often used to predict stock trends from its historical prices. Since there’s a lot of online information available in addition to stock prices, including technical indicators, news reports, and social information, we intend to combine news content for improving the performance. In this paper, we propose a deep fusion model for stock trend prediction combining news content with historical stock prices. Firstly, we utilize multi-layered Long Short-Term Memory (LSTM) to learn sequential information from stock prices. Then, we adopt Hybrid Attention Networks (HAN) which include both sentence-level and temporal attention to discover the relative importance of words from news reports. Finally, we compare early and late fusion models to improve stock trend prediction. The experimental results show that the best macro-F1 score of 79.0