
Massive online reviews play a crucial role in mining customer requirements. However, existing research has three limitations: (1) the inadequate capture of contextual semantics in traditional text classification models, (2) the lack of sentiment analysis targeted at the product attribute level, and (3) the neglect of inter-attribute dependencies in requirement analysis. To address these gaps, this study proposed a requirement extraction and analysis method (REAM) that integrated multi-sentiment reviews. Specifically, a multi-domain adaptive BERT–BiLSTM–hierarchical position-aware network (HPANet) was developed for the fine-grained classification of reviews to automatically categorize them into nine key product attributes (for example, cost-performance ratio and appearance) that served as concrete manifestations of customer requirements. For requirements analysis, this method examined three dimensions: customer attention to product features, sentiment orientation toward specific attributes, and inter-attribute dependencies, thereby facilitating a comprehensive understanding of customer requirements. In the requirement extraction phase, the experimental results demonstrated that the proposed model achieved 1.2% higher classification accuracy than the conventional BERT model on a custom-built customer requirement dataset. To further verify the stability of the BERT–BiLSTM–HPANet model, additional validation was conducted on public datasets (CNews and ChnSentiCorp), revealing significant improvements in both the accuracy and F1-score compared with the baseline models.
Financial fraud in listed companies has attracted increasing academic attention due to its severe implications for investors and market stability. Traditional methods for detecting fraudulent financial statements have proven insufficient in addressing the growing complexity and volume of financial data. This study proposes a novel hybrid model combining recursive feature elimination with cross-validation and parallel random forest (RFECV-PRF) for feature selection and a genetic algorithm-optimized LightGBM (GA-LightGBM) for financial fraud detection. The RFECV-PRF method effectively evaluates feature importance and selects the optimal subset of financial indicators, while the GA-LightGBM model enhances prediction accuracy and efficiency through global optimization of hyperparameters. Empirical tests using data from five industries—energy, materials, industrials, information technology, and healthcare—demonstrated significant improvements in classification accuracy, precision, recall, and F1 scores. The proposed framework outperforms traditional machine learning approaches such as random forest and XGBoost, achieving superior results in detecting fraudulent financial activities. This research highlights the potential of integrating advanced feature selection methods with optimized machine learning models to address the challenges of financial fraud detection and improve the reliability of corporate financial reporting.
A scientific and reasonable student evaluation system plays a crucial guiding role in the development of higher education. However, for the current student evaluation theories and models, their operability is relatively weak. The constructed models have poor adaptability (weak robustness), lack cross-scenario research on guiding policies, and the evaluation results of the models are also relatively single. Therefore, this study proposes a hierarchical student evaluation enhanced decision-making model which is based on a zero-order fuzzy classifier as the construction unit. By introducing the administrative decision-making characteristics of the educational administration department (such as policy orientation indicators, teaching intervention suggestions, etc.), combined with a hierarchical fuzzy system and an improved ridge regression algorithm, it achieves the collaborative optimization of evaluation efficiency and interpretability. The experimental results show that the model demonstrates excellent classification performance and semantic interpretability on the degree student evaluation dataset, and can accurately predict students’ academic performance and support personalized educational decisions.
English writing texts often show complex semantic layers, implicit emotional expression, and strong temporal dependence, which leads to limitations in semantic modeling and restricted feature extraction in existing methods. To address this issue, the study constructed a fine-grained sentiment recognition model that integrates bidirectional encoder representations from transformers (BERT), bidirectional gated recurrent unit (BiGRU), convolutional neural network (CNN), and an attention mechanism. BERT was used to generate context-aware semantic representations and improve overall semantic understanding of the text. BiGRU was applied to capture bidirectional temporal dependencies and describe the dynamic evolution of emotions in discourse. CNN was employed to extract phrase-level local emotional features and enhance the detection of emotion-triggering segments. The attention mechanism was introduced to highlight key emotional information and improve feature discriminability. On this basis, a gating fusion strategy was used to dynamically integrate multi-source features, and a multi-task learning framework was incorporated. The model performed emotion intensity prediction while conducting multi-class emotion classification. In this way, fine-grained sentiment modeling was achieved from both category and intensity perspectives. The results indicate that the proposed model achieves excellent performance across several datasets, with overall accuracy and F1-score remaining above 91%, ROC-AUC reaching up to 0.944, and recognition rates for all emotion categories staying above 84%, clearly outperforming existing mainstream approaches. Ablation results demonstrated that all six modules contributed to performance improvement. On the SST-2 and IMDB datasets, the complete model achieved an accuracy of 93% and an AUC of 0.97. In summary, this research offers a compact, stable, and adaptable solution for emotion modeling in English texts, with both theoretical and practical value.
As the emphasis on energy efficiency and environmental protection grows, an intelligent control strategy for the power generation process of gas boilers has become a key approach to optimizing industrial production. This paper presents an intelligent control strategy for a 150 MW ultra-high temperature subcritical gas boiler’s power generation process. The strategy aims to address the issue of frequent manual adjustments to the mixed gas and attemperating water valves due to fuel instability. The proposed strategy achieves stable control of the power generation process and dynamic regulation of the power generation load. It does so by monitoring key operating parameters of the gas boiler in real time, designing an intelligent control strategy that uses three key process variables (main steam temperature, gas equivalent, and attemperating water valve opening degree) as controlled parameters, and implementing expert rule-based control. Operational results demonstrate that this control strategy effectively stabilizes the main steam temperature within the specified process range and enables dynamic regulation of the power generation load. With a system utilization rate exceeding 90% and a reduction in standard coal consumption from 299.8 g/kWh under manual control to 297 g/kWh under automatic control, this strategy effectively stabilizes the power generation process. It significantly improves combustion efficiency, reduces energy consumption, and mitigates environmental pollution. Thus, it has promising practical application prospects.
The DEA-R model, which integrates data envelopment analysis (DEA) with ratio analysis, allows for the evaluation of efficiency using ratio data. One main strength of DEA is its ability to provide concrete improvement targets for inefficient decision-making units (DMUs). However, identifying these targets within the DEA-R framework is particularly challenging because of the inherent characteristics of ratio data. To address this issue, this study proposes a novel approach for identifying concrete improvement targets for inputs or outputs within the DEA-R framework. Specifically, we construct the DEA-R efficient frontier based on a unique concept and develop an approach to identify improvement targets that lie explicitly on this frontier. This ensures that all the identified targets are DEA-R efficient, thereby guaranteeing their rationality and validity. Furthermore, the constructed frontier enables the setting of flexible improvement targets under various scenarios, thereby enhancing the practicality and adaptability of the proposed approach.
This study investigates human impressions of autonomous agents that refrain from action. The increasing frequency of human–agent interactions has increased the demand for cooperative human–agent behavior. In human communication, cooperative behavior is frequently regarded as refraining from action to consider the requirements or desires of others (referred to as enryo in Japanese). However, enryo behavior is rarely investigated in human–agent interaction experiments. Investigating how people react to an agent refraining from action and observing actions of the agents would advance the society toward coexistence of humans and autonomous agents. Here, we demonstrate that humans indeed perceive enryo demonstrations in refraining-from-action agents; moreover, this behavior renders an agent more anthropomorphic and likeable than a non-considerate agent. Enryo impressions are also associated with increased human proactivity, suggesting that enryo functions as a social cue that encourages initiative actions by humans. Therefore, incorporating enryo -like behavior in agents may enhance the smoothness and cooperativeness of human–agent interactions.
In the era of globalization, learning English through oral means is becoming increasingly important. However, inaccurate speech recognition and imperfect feedback mechanisms have significantly hindered the improvement of the oral ability of learners. To solve this problem, this study proposes an English oral learning speech recognition and feedback system based on a multilayer improved long short-term memory network (MLSTM). The study uses the Texas Instruments and Massachusetts Institute of Technology (TIMIT) speech database and employs a hidden Markov model (HMM) and standard LSTM systems as controls to conduct a comprehensive test of the proposed model. The experimental results showed that the MLSTM model achieved high speech recognition accuracy, with an overall accuracy of 86.3% that was significantly higher than the 65.2% of HMM and 80.1% of LSTM. In terms of feedback information targeting and effectiveness, the MLSTM model scored 4.35 and 4.5, respectively, that were suggestively better than those of the comparison models. This showed that the MLSTM model could accurately recognize speech and provide learners with highly personalized and effective feedback. The research results enrich the theory of computer-assisted language learning, provide practical and effective tools for oral English learning, and promote oral English learning toward greater intelligence and efficiency.
Systematically comparing how linguistic representations relate to brain activity has become an important topic in the field of computational neuroscience. Prior studies have mainly relied on contextual hidden states combined with linear regression, leaving open questions about the role of static input embeddings and the benefits of nonlinear mappings. In this study, we compare input embeddings and hidden states from multiple language model families (BERT, GPT-2, and LLaMA) within both encoding frameworks, which map text features to brain responses, and decoding frameworks, which reconstruct linguistic features from brain activity. We benchmarked voxel-wise ridge regression against bidirectional long short-term memory (BiLSTMs) models, using repeat-split cross-validation and explainable variance normalization on functional magnetic resonance imaging (fMRI) data from three subjects. Our analyses demonstrate that input embeddings, despite being context-invariant, remain competitive and, in some cases, outperform hidden states, while BiLSTMs provide modest but region-specific improvements over ridge regression. Fine-grained voxel-level results further revealed distinct cortical distributions of stable versus context-dependent features. Together, these findings clarify the trade-off between predictive performance and interpretability and highlight that input embeddings offer a strong and interpretable baseline for representational alignment between language models and brain activity.
Aiming to address the problems of interest conflict between charging stations and electric vehicle (EV) owners, as well as severe load fluctuations caused by disorderly EV charging, this paper proposes a multi-objective optimal scheduling model based on an improved NSGA-III algorithm (TSM-NSGA-III). The model utilizes dynamic electricity price as a decision variable instead of a fixed time-of-use price, with optimization objectives set to maximize charging station profit, maximize EV owner satisfaction, and minimize the load peak-valley difference rate. The TSM-NSGA-III algorithm enhances the original NSGA-III through three key improvements: (1) chaotic reverse learning to improve initial population quality, (2) the sparrow search algorithm to avoid local optima, and (3) Manhattan distance to preserve population diversity and discover potential optimal solutions. Experimental results demonstrate that the proposed method achieves a 26% faster convergence and a 9.9% higher average solution quality compared to NSGA-III. Furthermore, it obtains superior Pareto frontiers with significantly better performance in both charging station revenue and user satisfaction, effectively overcoming the algorithm’s tendencies toward premature convergence and neglect of diverse optimal solutions.
Scene text recognition (STR) in natural images remains highly challenging due to the large variations in character appearance across diverse real-world conditions, such as changes in font, color, layout, and background complexity—which hinder model generalization and remain insufficiently explored. To address this issue, we propose a visual prompt-guided differential learning (VPDL) framework designed to improve the generalization capability of STR models without requiring scene-specific fine-tuning. Inspired by the human ability to reference prior visual knowledge when recognizing text, VPDL introduces a set of character-level visual prompts that guide the model in perceiving appearance variations among characters. Built upon these prompts, we develop a local-to-global differential learning strategy that enhances patch-level representations and aligns global features with character cues while preserving scene-specific information. Additionally, to mitigate exposure bias in autoregressive decoding, we replace conventional label inputs with context-aware textual prompts, encouraging the decoder to better utilize textual cues embedded in image features. Extensive experiments on widely used benchmarks and real-world datasets demonstrate the effectiveness of VPDL.
Precise orientation control of drilling tools is fundamental to directional drilling trajectory management. This study establishes a hybrid control framework for the high-accuracy regulation of inclination and azimuth. First, we derived a kinematic model characterizing the dynamic evolution of inclination and azimuth, explicitly addressing their coupled dynamics and azimuthal time-delay effects. A comprehensive downhole motion model was developed to capture system behavior. Distinct control strategies were formulated for the inclination and azimuth subsystems and integrated into a unified hybrid architecture. Stability analysis decomposes the system into individual control loops, with inclination stability ensured through conventional criteria, while azimuth stability is transformed into a solvable linear matrix inequality problem. Experimental validation demonstrates the superior robustness and engineering applicability of the method, achieving minimal attitude deviation under perturbed conditions. This control design provides a theoretically rigorous solution that bridges advanced control theory with practical well construction requirements.
The rapid advancement of digitalization is reshaping multiple aspects of firms and transforming the nature of innovation and entrepreneurship. However, a mature solution to accurately measure the digitalization maturity of enterprises is lacking. To help firms better understand their digitalization competitiveness, this study examined Chinese listed companies and, drawing on publicly available multi-dimensional heterogeneous data, employed the skyline algorithm and entropy weight method to construct a scientific measure and evaluation of digitalization maturity. To assess the reliability and practical feasibility of the measurement, this study further conducted heterogeneity analyses across industries and regions in China based on the evaluation results. The findings indicated that firms with superior digitalization maturity were predominantly concentrated in industries that were highly sensitive to digital technologies, as well as in regions characterized by stronger resource endowments and more frequent knowledge exchanges. In contrast, the digital transformation of traditional industries, such as the real estate sector, and of the regions where these industries are concentrated remained relatively weak and required further strengthening. These findings provide significant implications for both policymakers and industry stakeholders.
Recent advances in artificial intelligence and robotics have accelerated the deployment of service robots in daily environments. However, many systems still lack adaptive responsiveness to the nonverbal behaviors of users. This study proposes a gaze-based user analysis system integrated into a mixed reality (MR) smart home environment to support attentional intention-aware human–system interactions. Rather than directly estimating emotional states, the proposed approach infers attentional intentions of users based on gaze behavior as an operational proxy for emotional attunement. Using gaze data collected through HoloLens 2, we develop a machine learning model based on a long short-term memory network combined with a mixture density network to predict future gaze coordinates in a three-dimensional space. The predicted gaze information is shared with a robotic partner to enable proactive context-aware information support. The proposed system demonstrates the feasibility of leveraging gaze prediction to anticipate user focus and provide adaptive support in MR-based smart environments.
This study addresses the problems of imbalanced voice sequence, insufficient training stability, and symbol-audio modality mismatch in multi-track music generation. To this end, a gradient penalty-constrained multi-track adversarial training framework and a Wasserstein distance-driven cross-modal distribution alignment mechanism are studied and designed to optimize voice coordination accuracy and auditory perception authenticity. The core innovation lies in building a dual-module collaborative optimization architecture, pioneering gradient penalty constraints to eliminate multi-track training oscillations, and proposing Wasserstein feature space mapping mechanism to eradicate cross-modal perception mismatch, establishing a theoretical paradigm and technical path for joint optimization of voice and sound effects. Experimental verification: the voice conflict rate reaches 0.60%, the gradient norm variance is 1.13×10 -3 , and the training stability is controlled. The cross-modal distribution distance is 0.28, achieving precise alignment. The feature alignment error of 0.17 exceeds the technical limit, and the style fidelity is 91.7% (Baroque 94% / Jazz Blues 95%) to restore artistic expression. The dynamic expressive power is 4.66 points, approaching human creativity, and the auditory similarity is 0.89, establishing perceptual authenticity. The conflict rate of voice parts in the ablation test decreases by 52%, and cross-modal mismatch compression is reduced by 49%. Parameter sensitivity analysis shows that when the hidden space dimension is 64, the structural entropy is 0.810 and the perceptual similarity is 0.87, reaching the global optimum. This model significantly improves the coordination of multi-track structures and cross-modal perception quality, providing core technical support solutions for industrial-grade artificial intelligence (AI) music creation platforms such as film and television music composition and digital composition.
This study briefly introduces an intelligent detection algorithm for foreign objects on solar panel surfaces, as well as an intelligent cleaning robot. In the intelligent detection algorithm, the improved Retinex algorithm was used to improve low-light images, and the You Only Look Once version 5 (YOLOv5) algorithm was used to detect foreign objects on the surface. Simulation experiments were performed. The improved Retinex algorithm was compared with the traditional Retinex and histogram equalization methods. The YOLOv5 algorithm was compared with the faster region-based convolutional neural network (R-CNN) and YOLOv4 algorithms. The surface foreign object cleaning ability of the developed intelligent robot was compared with the robot that did not use the same algorithm. The results showed that the improved Retinex algorithm could increase image brightness while preserving color. The edge strength, information entropy, and locally orderless error of the improved images were 79.8±1.6, 7.5±0.7, and 813.6±2.6, respectively. The YOLOv5 algorithm could identify and locate foreign objects more accurately, with a precision of 0.987, a recall rate of 0.985, and an F -value of 0.986. It was also discovered that the intelligent robot using the proposed surface foreign object detection algorithm cleaned foreign objects on the surface of photovoltaic panels faster and better. The time consumed in one round of cleaning was 6.2±0.1 min, and the residual foreign object on the surface was 0.7%±0.1%.
This paper proposes a category-centric initialization that introduces prior knowledge for knowledge graph embedding (KGE) at a low cost. KGE is a technology that maps symbols to embeddings to utilize large-scale knowledge graphs, and it has been widely applied because of its simplicity and efficiency. However, the initialization challenge with this technology has long been overlooked. This critical issue has implications for the training cost, the stability of training, and even the final performance of the model. KGE predominantly utilizes random initialization, which overlooks the wealth of prior knowledge embedded within knowledge graphs. To counteract this, pre-training initialization has been introduced as a way to utilize the prior knowledge. While this strategy can lead to enhanced model performance and quicker convergence rates, it increases computational demands and restricts application breadth. To address these challenges, we propose a novel initialization called category-centric initialization (CCI). CCI is designed to be universally applicable across any scenario involving the training of KGE models from scratch. It utilizes the weighted sum of the category embedding and the random embedding as the initial embedding of entities. By integrating explicit category information into the random initialization, CCI effectively utilizes prior knowledge while avoiding excessive computational cost. The results of experiments demonstrate that the proposed method can effectively reduce the training cost of advanced KGE models without degrading the final performance. Additionally, the results of experiments without category information show that our method can be applied in scenarios where explicit categories are not given to entities.
An electro-hydraulic drive system is essential for the stable operation of tunnel drilling rigs in underground coal mines. However, components such as pumps, valves, and controllers inevitably experience gradual degradation under long-term and high-load conditions. Conventional monitoring approaches often rely on labeled fault data or suffer from limited interpretability, restricting their applicability in real engineering environments. To overcome these limitations, this study proposes an unsupervised degradation trend analysis method that does not use labeled samples. A sliding-window strategy was adopted to extract key statistical features. Principal component analysis was then employed to construct a unified health index, and Z-score normalization enabled the interpretable detection of abnormal tendencies in individual features. Validation on real drilling data revealed clear degradation behaviors, such as main pump leakage and control current drift, demonstrating that the proposed method offered a lightweight and interpretable solution for trend-based condition monitoring and provided practical support for the intelligent maintenance of electro-hydraulic drive systems.
Pre-trained language models (PLMs) have demonstrated high performance across various tasks and domains. Among these PLMs, Mixture of Experts (MoE) models also exhibit high performance with fewer active parameters. In domain adaptation, generally, continual pre-training is performed with existing models using domain-specific corpora. However, few efforts have been made to transform and train these models into models with MoE architectures. We propose a method to construct domain-adapted MoE models from general pre-trained models that do not initially have MoE architectures. By independently training multiple experts using domain corpora and integrating them into an MoE architecture, we constructed a domain-adapted MoE model. We performed this MoE transformation in the financial domain and verified its effectiveness in financial tasks. The evaluation results indicate that our domain-adapted MoE models perform better than those without MoE architectures. Our domain-adapted MoE models achieved consistent improvements over domain-adaptive pretraining, task-adaptive pretraining, and domain- and task-adaptive pretraining, with an average improvement of approximately 0.08 in F1 across six financial benchmark tasks for encoder-based models.
Accurate forecasting of the hearth lining temperature in blast furnaces is essential for operational safety and efficiency, however it remains challenging owing to the complex spatiotemporal coupling and time-lag effects among process variables. To address this issue, we present a new spatiotemporal feature modeling framework that integrates gated recurrent units (GRUs) with a dual-attention mechanism to capture multi-scale temporal dependencies and dynamically assess variable importance. A convolutional neural network module is incorporated to extract localized spatial features from the time-series data, thereby enhancing the representation of the underlying metallurgical mechanisms. The validation on real industrial data showed that the proposed model achieved a root mean square error of 0.0523 and a hit rate of 94.26%, outperforming conventional long short-term memory and GRU models. This approach offers a reliable solution for intelligent health monitoring and proactive maintenance in modern data-driven ironmaking operations under highly dynamic and uncertain conditions.