
In the area of image recognition, the quality of image label data has a significant impact on the performance of classification models. Therefore, manual annotation has been used as a means to label images. However, manual annotation is laborious and time-consuming and can introduce additional noise. To address these issues, this paper investigates an automatic algorithm for improving a cat breed classification model based on meta loss correction. The proposed algorithm leverages web crawling techniques to obtain unlabeled images of cats, filters them through object recognition, and selects only images containing cats. These images are then fed into the algorithm, which utilizes a pretrained initial model to generate pseudo-labels. These pseudo-labeled data are subsequently refined using a meta loss function, correcting the inaccuracies associated with the pseudo-labels. Finally, the labeled new data is merged with the original dataset, gradually increasing both the quantity and quality of the dataset. Experimental results demonstrate that as the merged dataset expands, the model's error decreases gradually, and its performance improves.
Purpose This paper aims to disentangle Chinese-English-rich resources linguistic and speaker timbre features, achieving cross-lingual speaker transfer for Cambodian. Design/methodology/approach This study introduces a novel approach: the construction of a cross-lingual feature disentangler coupled with the integration of time-frequency attention adaptive normalization to proficiently convert Cambodian speaker timbre into Chinese-English without altering the underlying Cambodian speech content. Findings Considering the limited availability of multi-speaker corpora in Cambodia, conventional methods have demonstrated subpar performance in Cambodian speaker voice transfer. Originality/value The originality of this study lies in the effectiveness of the disentanglement process and precise control over speaker timbre feature transfer.
Mobile crowdsourcing (MCS) can solve problems that are difficult for computers to solve accurately or efficiently. Current crowdsourcing workers face the challenges of overload, task recommendation is presented to deal with the above issue. The existing MCS task recommendation methods only consider the workers themselves, without taking the potential relationships between workers and workers and tasks, which may fail to grasp the long-term behavioral preferences of workers and cannot solve the cold start problem efficiently. To solve the above problems, this paper proposes a task recommendation method based on heterogeneous graph link prediction. First, we try to obtain the connection relationships between crowdsourcing tasks and workers and build a worker-task heterogeneous graph accordingly. Then, majority voting and DeepWalk are used to generate initial node features of workers. Second, we design two efficient message passing mechanisms to aggregate and update node features between task and worker nodes to explore the potential relationships between nodes. Finally, we obtain attention weights among nodes and use Bi-GRU to capture the long-term behavioral preferences of workers and recommend appropriate tasks to workers based on the similarity between workers and tasks to improve perceptual quality. Evaluations are conducted on seven real datasets. Experimental results show that our method is superior to the state-of-the-art baselines.
Microservice workflows are widely used in real-time mobile computing scenarios such as face recognition and speech recognition. The key challenge is to develop efficient, stable, and robust algorithms capable of handling uncertain and fuzzy workflow tasks. In this paper, we consider the real-time microservice workflow fuzzy scheduling problem under VM-Container two-tier resources. A novel model based on triangle fuzzy numbers is formulated. The model encompasses two metrics: cost and degree of satisfaction. In response, this paper proposes an ARIMA prediction-based workflow fuzzy scheduling method (PFSM), comprising five key components: prediction of workflow arrival number, predistribution of task sub-deadlines, ordering of the task pool, task scheduling strategy, and resource management. To assess the performance of the proposed algorithms, several comparison algorithms are selected for analysis, and their performance differences are evaluated using ANOVA. The experimental results demonstrate the significant superiority of the proposed algorithms over the other compared algorithms in terms of overall performance.
Knowledge tracing is the fundamental technology for constructing learner models that dynamically estimate and predict a learner's knowledge state. While current research on knowledge tracing has improved the predictive capacity of the model by investigating the relationship between learners and problem concepts, these models become static after training. This limits their ability to adapt to the varying developmental stages of learners due to human diversity. Drawing inspiration from the synaptic plasticity that grants lifelong learning capabilities to the biological brain, this study incorporates plasticity weights into the Transformer architecture. This leads to the proposal of a deep knowledge tracing model designed to adapt to individual learner development. Moreover, it demonstrates significant performance improvements compared to the baseline model.
Multi-robots collaborative has important application prospects in fields such as military, medical, transportation, security patrols, rescue and disaster relief. Robots keep formation and obstacle avoidance are two of the core technologies in the application of multi robot cooperative system. This paper designs and implements a distributed robots formation and obstacle avoidance system based on the combination of fuzzy cascade PID and improved artificial potential field. This paper also designs a model of 2D lidar used to judge obstacles. It ensures the stability of multi robot formation through cascade fuzzy PID, When multi-robots system encountering obstacles, Combine with the cascade fuzzy PID and the improved artificial potential field algorithm, the multi-robots can complete obstacle avoidance while ensuring the formation. Finally, experiments were conducted on multi-robots physical platform, and the results showed that the algorithm can ensure formation during robot movement and avoid obstacles when moving in the scenario with massive obstacles.
With the development of transportation, the traditional traffic signal systems being unable to provide dynamic and flexible timing schemes for urban arterial road traffic in complex lane queuing conditions. In the control of arterial traffic, to solve the problem that vehicles queuing in turning lanes of branch road and then congesting the arterial road, this paper proposed an arterial traffic optimization algorithm based on deep reinforcement learning (DRL) and green wave coordination control in complex lane queuing conditions. The proposed algorithm provides a detailed division of the arterial roads and analyzed the mutual influence between vehicles inside the roads, combines DRL algorithm with the MAXBAND algorithm to optimize the signal period, phase sequence and green signal ratio of arterial roads, creates a new reward function for Deep Q Network (DQN) algorithm for multi-agent coordination. The algorithm was validated in SUMO simulation environment. The simulation results prove that the algorithm can flexibly perform signal timing and is more effective than traditional algorithms.
Existing crowdsourcing research has traditionally focused on high-skilled workers during the task allocation phase while neglected the potential value and skill enhancement opportunities of workers who are with lower skill levels. Consequently, as the number of completed tasks increases, skill disparities among workers have worsened. This can be primarily attributed to the platform’s heavy reliance on high-skilled workers, which makes it difficult for lower-skilled ordinary workers to meet task requirements and limits their chances for skill development. As a result, when a large volume of tasks awaits allocation, the limited number of high-skilled workers falls short of meeting the task demand, resulting in prolonged task completion periods and reduced task allocation success rates. In reality, the majority of tasks on crowdsourcing platforms contain straightforward components that can be handled by ordinary workers. Additionally, collaborative efforts among multiple workers can enhance task execution efficiency and reduce task durations. Building on these practical observations, this paper proposes a task allocation model and simulates the improvement of worker skill levels. In situations where teams take a longer time to complete tasks, the model allows ordinary workers to join the team, thereby assisting professional workers in enhancing work efficiency. Experimental results demonstrate that this model can alleviate worker skill disparities, diminish task durations, and facilitate the development of crowdsourcing platforms.
Speech-driven 3D facial animation is still an intensive field of research, with some persistent challenges. These difficulties arise from the intricate nature of achieving facial realism and the scarcity of audiovisual data. Previous studies have mainly focused on learning phoneme-level features from brief audio segments, which often lead to suboptimal lip movements. To capture the nuances of facial expressions, such as eyebrow-raising or lip curling, our proposed solution builds upon the autoregressive model of the Transformer, which can generate realistic facial movements based on previous frames and introduces a modular face separation model, which can separately control the upper face and lips, to enhance the quality of voice-driven 3D facial animation. This novel modular separation technique divides the facial mesh into two parts: the upper face and the lips, using our uniquely designed mask. Such an approach not only significantly improves facial animation synthesis but also lays the foundation for future research and application in this domain.
As a vital methodology for collaboration problems, Role-Based Collaboration (RBC) consists of three essential stages: agent evaluation, Group Role Assignment (GRA), and role transfer. Among these stages, GRA holds significant importance as a critical step in the process. In Social Sciences, trust between group members in collaboration tends to profoundly impact the final performance. However, existing GRA problems neglected the trust relationship between agents, which plays a decisive role in whether agents can work together. Therefore, this paper discusses the GRA problem with the constraint of trust between agents, aiming to maximize group performance while satisfying the trust constraint. The contributions of this paper include: 1) A new problem of GRA called GRA with trust between agents (GRATA) is formalized; 2) A feasible solution and a theorem are proposed to address the GRATA problem effectively; 3) Comprehensive experiments are conducted to validate the benefits of solving the problem.
Traffic signal control plays an important role in reducing urban traffic congestion. In complex traffic scenarios, coordinating phase signal control between intersections is a significant challenge. Reinforcement learning is widely used in the field of intelligent traffic signal control because it is good at dealing with sequence decision problems. The current reinforcement learning based approach makes phase decisions through coordinated cooperation. However, existing methods have difficulty with information exchange, because they lack semantic interpretation and explicit quantification of collaborative impact, which results in inefficient or conflicting phase coordination between intersections. Moreover, during the early exploration stage of reinforcement learning training, the phase output of the decision network is unreliable, making it difficult for the model to utilize decision information for self-supervised training. To address these issues, this paper proposes a self-supervised, explicit coordination based multi-agent reinforcement learning approach. Additionally, a phase boosting learning from demonstration method is introduced in the early training stages. Extensive experimental results demonstrate that this method can enhance collaboration among agents, outperforming baseline methods across multiple real-world traffic datasets, while also improving training stability and convergence speed.
In response to the limited and fixed scheduling patterns in existing intelligent scheduling technologies, this paper proposes a Memetic algorithm with Exchange Coding (MA-EC) based on flexible and non-uniform work schedules to develop rational scheduling plans that meet the needs of employees and production plans, thereby improving employee satisfaction and enhancing enterprise competitiveness. Firstly, three-dimensional binary coding is used to clearly express population individuals. Secondly, a greedy initialization of population individuals is performed to meet work requirements. Thirdly, the exchange coding method is utilized to update the population and reduce the search space, accelerating the convergence rate of the algorithm. Finally, a local search based on employee preferences is designed to avoid the algorithm from getting trapped in local optima and enhance the global search ability of the algorithm. Experimental results on instances with nine problem scales generated randomly show that compared with competing algorithms, the proposed algorithm has faster convergence speed, higher search efficiency, and can obtain intelligent scheduling plans with higher employee satisfaction.
Temporal sentence grounding in video (TSGV) focuses on identifying the most pertinent temporal segment within an untrimmed video, given a natural language query. Its principal aim is to ascertain and retrieve the specific moment that impeccably aligns with the given query. Although the existing methods have done much research in this field and achieved specific achievements, there are still problems of massive calculation and insufficient grounding. Our method mainly focuses on obtaining better video and query features performing cross-modal feature fusion better, and locating more accurately when dealing with this problem. We propose an efficient Cross-Modal Grounding Network (CMGN) to balance the amount of computation and localization accuracy. In our proposed structure, we obtain the local context information through a bidirectional Gated Recurrent Unit (GRU). We obtain the start and end boundary characteristics for a better video presentation. Then, the two-channel structure, divided into a start channel and an end channel, captures the temporal relationships among several video segments sharing common boundaries. To validate the effectiveness of our method, extensive experiments were conducted on two datasets.
Multi-agent patrolling has significant implications for addressing real-world security concerns. In multi-agent systems, the actions of agent directly influence those with whom it interacts. Traditional reinforcement learning-based multi-agent security patrolling methods overlook the role of these localized interactions in coordination among agents, thus failing to enhance the efficiency. To address this issue, this paper introduces a security patrolling approach based on Structured Coordinated Proximal Policy Optimization (PPO). The multi-agent patrolling task is modeled as a finite-time-step distributed partially observable semi-Markov decision process. This method, grounded in the Shapley Value, designs a multi-agent credit allocation function. The efficiency of this function is amplified using the structure of localized interactions. By accurately evaluating the contributions of each agent's selected actions, this function fosters enhanced coordination among agents. Extensive experiments in various scenarios were conducted, and the results demonstrate that our algorithm outperforms benchmark algorithms in terms of convergence speed and patrolling performance.
Text review is a task that determines whether the knowledge expression in a student answer is consistent with a given reference answer. In the professional scenarios, the number of labeled samples is limited, usually ranging from dozens to hundreds, which makes the text review task more challenging. This paper proposes a text review method based on data augmentation, which is performed by the combination of different positive and negative labeled samples. The review model infers the unlabeled samples, where the pseudo-labeled samples with the high confidences are selected for the subsequent training rounds. Experimental results in real national qualification exam datasets show that our method has improvement compared with the traditional method on the text review task under the limited sampling constraints.
The large-scale optimization problems (LSOPs) have been a hot research in evolutionary computation (EC) community. Although there have been many contributions from researchers in solving LSOPs, the large search space and the numerous local optimal solutions of LSOPs are still two important challenges. In order to alleviate the above challenges, this paper proposes a dimension-based elite learning particle swarm optimizer (DELPSO). In DELPSO, individuals in the population have their unique update probabilities and select specific learning exemplars according to their own properties, making the evolution of the population more efficient. Meanwhile, in the evolutionary process, each individual chooses two different learning exemplars for each dimension, so that each individual can learn from multiple learning exemplars and using the information from multiple individuals to help its own evolution and enhance the diversity. To testify the effectiveness of the proposed algorithm, DELPSO and some large-scale algorithms are experimented on a widely used large-scale benchmark suite IEEE 2013 and the experimental results show that DELPSO outperforms other comparative algorithms in general.
Low accuracy and wear are common issues with single photoelectric encoders in track positioning. A deep learning-based method for clip detection, counting, and positioning correction can effectively improve the accuracy of the photoelectric encoder. In this paper, we propose a subway track clip detection model based on MobileNetV3-YOLOv5s, a clip counting and positioning model based on DeepSort, and a fusion correction model for positioning data. Finally, through comparative experiments, we validate that our adopted method achieves higher positioning accuracy and stronger reliability.
It is very important to properly handle public emergencies, such as accident disasters, public health events and social security events, and understanding public opinion on public emergencies and its evolution is necessary to deal with the public emergencies. In this paper, we focus on nine public health events related to COVID-19, and explore the evolution of public opinion on these events. Specifically, we first collect information of public opinion on an event from Sina Weibo, including posts, comments. Based on the collected data, commenting networks are constructed. After that, we design a method to explore the evolution of public opinion on these events by observing and analyzing the evolution of commenting networks, including the changes in the number, emotions and topics of the comments. Further, we analyze the influence of emotion on the number and the topics of the comments. Finally, we obtain some observations that can help the emergency management departments understand the evolution of public opinion on public health events, and developing emergency plans to guide and control it.
The underwater communication environment is complex and changeable, with special requirements for delay and packet loss rate. This paper proposes a cluster-based cooperative sensor network with switchable topology, MAC layer protocol, and routing protocol for underwater environments. Through the dynamic selection of cluster heads and cluster members, and the reasonable setting of the topology structure, the sensors in the network can cooperate in communication and aggregate information to the sink node. The protocol and cluster head nodes are dynamically adjusted according to the network conditions, and the MAC protocol and routing protocol are dynamically switched to ensure that all nodes in the network can successfully cooperate in communication. At the same time, nodes can switch states to ensure the robustness and network life of the entire network.
Automatic and accurate instance segmentation of teeth from 3D Cone-Beam Computer Tomography (CBCT) images is crucial for dental diagnose. Although Convolutional Neural Networks (CNNs) are widely used for tooth instance segmentation, the limitations of CNNs in capturing global image information can impact model performance. Recently, Transformer models leveraging the Self-Attention mechanism have exhibited exceptional capabilities in modeling global relationships in images. In this paper, we propose a fully automated tooth instance segmentation model utilizing the Self-Attention mechanism. The model is primarily based on the Self-Attention UNETR++ network and consists of three stages. In the first stage, a V-Net is employed to identify the region of interest (ROI) containing the teeth. In the second stage, a multitask UNETR++ network is utilized to extract the centroid and skeleton of the teeth. In the third stage, another multitask UNETR++ is employed to simultaneously learn the tooth mask and boundary, leading to accurate tooth instance segmentation. Experimental results on a dataset consisting of 98 CBCT images demonstrate the efficacy of our method. It achieves a Dice score of 95.1 % and reduces the average surface distance (ASD) to 0.14 mm.