With the rapid development of big data and artificial intelligence technologies, large language models (LLMs) are increasingly being applied across multiple fields. In the healthcare domain, efficient utilization of extensive patient data and medical records considerably enhances diagnostic accuracy and enables disease-risk prediction. This study aims to establish a research framework for evaluating the application of assessment models in structured and unstructured data, exploring the potential applications of LLMs in healthcare. This study proposes the HLLM-Potential Framework (Healthcare Large Language Model Potential Evaluation Framework), a comprehensive evaluation framework designed to assess the applicability and performance of LLMs in medical data analysis through comparative experiments with traditional models. Publicly available and standardized cardiovascular datasets were adopted, covering both structured and unstructured data. Existing LLMs were utilized through training and task-specific configuration on medical data to perform disease prediction and health-risk assessment. In addition, various LLMs were systematically compared with traditional machine-learning and deep-learning models to quantify the differences in their predictive performance. Existing LLMs can process structured and unstructured medical data and use them to predict diseases and evaluate health risks. For structured cardiovascular disease prediction tasks focusing on heart failure, all considered LLMs achieved accuracy and recall rates above 80
Financial statement fraud detection (FSFD) is crucial for sustaining capital market stability, protecting investor confidence, and reducing systemic risks. Conventional approaches often rely on resampling or minority class sample synthesis techniques to mitigate class imbalance in FSFD. However, these methods face challenges due to poor interpretability of synthetic samples and inadequately capturing distributional shifts between fraudulent and normal samples while decoupling data balancing from classifier training. To overcome these limitations, a joint optimization framework for FSFD is proposed. Firstly, this framework constructs the GAN generator as a reinforcement learning-based "Counterfeiter" to transform the normal samples into fraudulent ones. Then, the generator is trained using a multi-objective reward function that simultaneously enhances sample fidelity and domain consistency. Finally, a closed-loop feedback mechanism dynamically evaluates synthetic samples and integrates them into classifier learning process, thereby achieving end-to-end training. Experimental studies on real financial data from the Chinese stock market demonstrate the superior performance of both the co-optimized classifier and the generated samples compared to baseline methods, with an improvement of 4.40% in F1-score, 4.44% in MCC, and 0.37% in AUC. Ablation studies further validate the necessity of each key module. Moreover, we introduce a novel interpretability analysis method that statistically characterizes perturbation magnitudes during state transitions, revealing potential risk indicators for financial manipulation.
Credit bonds play a vital role in financing economic development but also threaten financial stability through default risks. Current research primarily focuses on bond characteristics while overlooking structured knowledge of inter-bond relationships. Although GNN demonstrates advantages in modelling relational knowledge, their performance is substantially constrained by class imbalance in credit default prediction. To address these challenges, we have constructed a new knowledge-infused credit bond network dataset, which explicitly encodes real-world bond correlation information derived from market data. Building upon this dataset, we propose LANR-GNN, a novel graph neural network that differs from prior GNN-based approaches by integrating feature evaluation, label-aware neighborhood resampling, feature concatenation, and a minority-class weighted loss to better capture relational dependencies under imbalance. Empirical results demonstrate that LANR-GNN effectively exploits bond relationship knowledge and outperforms baselines in handling imbalanced relational data. Additionally, the model exhibits generalization capability on public datasets. Our work provides financial regulators with a novel model for systemic risk monitoring and offers methodological insights for imbalanced graph learning.
Anomaly detection under positive-unlabeled (PU) learning remains challenging as only a few anomalies are labeled, whereas the majority of data remain unlabeled and maybe contaminated by unknown anomalies. Existing PU learning approaches often rely on imprecise class-prior estimation or biased synthetic sample generation, potentially leading to unstable optimization and degraded detection performance. To mitigate these limitations, we propose PUGANomaly, a novel PU learning method for anomaly detection based on dual-branch generative adversarial networks (GANs). PUGANomaly integrates anomaly detection and PU learning within a unified architecture. By incorporating a self-training mechanism and a joint anomaly scoring strategy, PUGANomaly iteratively refines pseudo-negative samples and model parameters from anomaly similarity and normal dissimilarity perspectives. Furthermore, we design a set of PU-specific loss functions to explicitly enforce increased separation between normal and anomalous representations. To further strengthen discriminative capability, an auxiliary classifier is incorporated to learn the decision boundary between labeled anomalies and pseudo-normal samples, and the outputs are subsequently fused with reconstruction-based anomaly scores through an ensemble learning module. Extensive experiments conducted on the MNIST, FashionMNIST, and Alzheimer datasets demonstrate that PUGANomaly consistently and statistically outperforms state-of-the-art GAN-based PU learning models on anomaly detection tasks, achieving superior detection accuracy and robustness under limited supervision.
With the intensification of global population aging, the incidence of cognitive disorders such as dementia continues to rise. The Mini-Mental State Examination (MMSE) and other alternative tools can help doctors detect subtle changes in cognitive function at an early stage. These assessment tools can make a diagnosis before symptoms become severe, providing opportunities for early intervention, which is crucial for delaying disease progression and improving the quality of life of patients. However, traditional cognitive assessment methods are overly complex and affected by various factors. With the development of artificial intelligence technology, many new assessment tools are constantly being developed and improved. How to evaluate the effectiveness of intelligent electronic cognitive assessment tools is particularly important. We have proposed the Correlation and Supervised Learning-based Cognitive Tool Effectiveness Assessment Method (CSL-CTEA) to evaluate the effectiveness of intelligent electronic cognitive assessment tools, including: (1) experimental design and data collection based on traditional scales and intelligent electronic assessment tools, (2) consistency and correlation tests; (3) accuracy analysis of assessment results based on supervised learning. We used CSL-CTEA to explore the effectiveness of a certain electronic assessment. This intelligent electronic cognitive assessment tool includes voice tests, orientation tests, and picture recognition tests to assess cognitive abilities from multiple perspectives. The results show that the electronic assessment is in good agreement with traditional cognitive assessment methods. The various indicators of the electronic assessment can explain the changes in MMSE scores to some extent. The study also found that the electronic assessment performs well in determining whether the subject is at cognitive risk. To some extent, the electronic assessment can replace traditional cognitive assessment methods such as MMSE to help people judge whether they are at risk of cognitive decline.
With the intensification of global population aging, the incidence of cognitive impairment such as dementia continues to rise. The Activity of Daily Living (ADL) scale can help to assess daily living functions and early screen for dementia, which is crucial for delaying disease progression and improving the quality of life in older adults. As advanced ADL assessment tools continue to be developed and improved, how to evaluate their effectiveness is particularly important. However, most studies have assessed these tools from a single perspective and have often failed to examine the contribution of individual items within the scales. Therefore, we propose the Statistical and Machine Learning-based Advanced ADL scale Effectiveness Assessment Method (SML-AAEA) to evaluate the psychometric properties and early dementia screening ability of ADL assessment tool, including: (1) scale design and data collection based on traditional scales and advanced items; (2) analysis of scale validity and reliability; and (3) analysis of scale dementia diagnostic ability and item importance using machine learning. We then apply SML-AAEA to investigate the effectiveness of our proposed Advanced ADL scale for Early Dementia Screening (AADLs-EDS), which introduces three new advanced items, namely “Going far away”, “Online shopping” and “Using smartphone”. The results show that AADLs-EDS has excellent construct validity, measurement invariance, and scale reliability. The total score of AADLs-EDS can explain the changes in elderly cognitive functions to some extent. The study also finds that AADLs-EDS outperforms the traditional ADL scale in classifying dementia, with the three new items showing the strongest predictive contributions. The findings confirm that AADLs-EDS is a reliable and valid tool for early dementia screening.
Semiconductor chip is the core of the digital economy, has important strategic significance, in recent years, the global semiconductor market size is growing rapidly. At present, the domestic semiconductor industry is still in the late stage relative to the leading enterprises with first-mover advantages, and is in urgent need of rapid and stable development. The international and domestic semiconductor market has been analysed and focuses on the opportunities and challenges currently facing the domestic semiconductor industry. Based on the semiconductor component risk assessment data, the risk value of semiconductor components is modelled and analysed from a microscopic point of view using the CART regression tree method in data mining, and the model accuracy is further verified and improved by XGBoost, which analyses the microscopic influencing factors of semiconductor component risk, and is an important revelation for the construction of a complete semiconductor component risk prediction mechanism in China.
Background The increasing aging population has led to a shortage of geriatric chronic disease caregiver, resulting in inadequate care for elderly people. In this global context, many older people rely on nonprofessional family care. The credibility of existing health websites cannot meet the needs of care. Specialized health knowledge bases such as SNOMED—CT and UMLS are also difficult for nonprofessionals to use. Furthermore, professional caregiver in elderly care institutions also face difficulty caring for multiple elderly people at the same time and working handovers. As a solution, we propose a smart care system for the elderly based on a knowledge graph. Method First, we worked with professional caregivers to design a structured questionnaire to collect more than 100 pieces of care-related information for the elderly. Then, in the proposed system, personal information, smart device data, medical knowledge, and nursing knowledge are collected and organized into a dynamic knowledge graph. The system offers report generation, question answering, risk identification and data updating services. To evaluate the effectiveness of the system, we use the expert evaluation method to score the user experience. Results The results of the study showed that compared to existing tools (health websites, archives and expert team consultation), the system achieved a score of 8 or more for basic information, health support and Dietary information. Some secondary evaluation indicators reached 9 and 10 points. This finding suggested that the system is superior to existing tools. We also present a case study to help the reader understand the role of the system. Conclusion The smart care system provide personalized care guidelines for nonprofessional caregivers. It also makes the job easier for institutional caregivers. In addition, the system provides great convenience for work handover.
The identification of disruptive technologies is paramount for the expansion into novel markets and the creation of innovative value paradigms. Despite the importance of this task, current methodologies for identifying such technologies are predominantly qualitative, which are often hindered by a dearth of empirical data and a predisposition towards subjectivity. This paper proposes a paradigm shift by employing patent text data and utilizing Latent Dirichlet Allocation (LDA) for data preprocessing to extract and analyze technical themes from patent abstracts. The research innovatively posits the problem of identifying disruptive technologies as an anomaly detection issue and addresses it through the application of a one-class Support Vector Machine (SVM) algorithm. In the empirical component of the study, the field of gene editing is selected as a case study to validate the proposed method’s efficacy and practicality. The validation is conducted through a multifaceted analytical framework, which includes: an examination of the distribution and proportion of patent topics, a comprehensive market analysis, and a comparative analysis of the proposed model against existing methodologies. The results of this study not only substantiate the potential of the proposed approach in accurately detecting technologies with disruptive potential but also highlight its robustness across various analytical dimensions.
The Activity of Daily Living Scale (ADLs) has been extensively used to evaluate the fundamental activity of daily living (ADL) of the elderly. However, with societal advancements and improvements in living standards, people’s life behaviors and activity abilities change, resulting in traditional ADLs potentially inadequate for assessing contemporary ADL. This paper proposes an enhanced version of the traditional ADLs, incorporating two key aspects, basic mobility improvement and smart device application. Three items are introduced, namely “Going far away”, “Using smartphone”, and “Online shopping”. An experimental design and sample study were conducted to evaluate the reliability, validity, and accuracy of the new scale using appropriate methods. The findings indicate that the new ADLs (NADLs) has certain advantages in assessment content, reliability and validity, decline effect and accuracy, compared with the traditional ADLs. Specifically, the improved scale assesses the higher level and more complex ADL of the elderly, with the inclusion of supplementary items. Furthermore, the NADLs demonstrate superior accuracy in assessing ADL among the elderly. The NADLs also exhibit a lower ceiling effect compared to the traditional ADLs, indicating better content effectiveness. In conclusion, with societal progression, the NADLs emerges as a reliable and valid measure, serving as an effective tool for evaluating ADL among the elderly.
This study aims to improve the accuracy of detecting financial fraud in listed companies by applying various deep learning algorithms. First, we comprehensively reviewed the current state of research on financial fraud theory and identified 67 recognition features to create a new recognition indicator system. Then, we collected data samples of all A-share listed companies from 2010 to 2022, preprocessed them, and generated a basic dataset. To address the unbalanced dataset, we used the Borderline-SMOTE algorithm. Empirical analysis results show that this algorithm can significantly improve the recognition performance of the model. Finally, we conducted experiments on the new dataset using three types of deep learning algorithms. The results show that the model constructed using the Long Short-Term Memory (LSTM) algorithm has the best prediction performance, with an accuracy rate higher than that of the DCRN, autoencoder, and other models. In addition, the classification effects of all deep learning algorithms are better than basic models and ensemble models. This research provides a powerful tool for the regulatory authorities of listed companies, helping them more effectively monitor and prevent financial fraud. We have three innovations in this study: (1) Development of a comprehensive recognition indicator system with 67 features; (2) Utilization of the Borderline-SMOTE algorithm to handle data imbalance; (3) Demonstration of the superior performance of the LSTM algorithm compared to other deep learning, basic, and ensemble models.
The valuation and pricing of data assets is not only a key factor in the national strategic development but also of extreme importance for corporate operations. This article proposes to explore the valuation and pricing of data assets from the perspective of intangible assets based on the theory that “data assets are a component of intangible assets.” Secondly, it focuses on the internet platform companies that rely on massive user data. Considering that a company’s value is influenced by various asset values (including the value of fixed assets, intangible assets, current assets, etc.), this article constructs an input-output indicator system based on corporate assets, and describes the value of data assets from three aspects: data quality, data application, and data risk. Ultimately, the total value of a company’s data assets is obtained through the proportion of various data asset indicators in the company’s total market value.
INTRODUCTION:Effective behavioral management is critical for people with diabetes to achieve glycemic control. Many instruments have been developed to measure diabetes-specific self-management. This review aimed to retrieve existing self-management-related instruments and identify well-validated instruments suitable for clinical research and practice. METHODS:First, PubMed, Psych INFO, ERIC, and two Chinese databases (CNKI and Wanfang Data) were searched to identify existing instruments for self-management in diabetes systematically. Second, instruments were screened based on the pre-specified inclusion and exclusion criteria. Third, the psychometric property data of each included instrument were retrieved, and instruments with poor psychometric properties were excluded. Fourth, selected instruments were categorized into four categories: knowledge and health literacy, belief and self-efficacy, self-management behaviors, and composite scales. Finally, recommendations were made according to the application status and quality of the instruments. Instruments in English and Chinese were screened and summarized separately. RESULTS:A total of 406 instruments (339 English instruments and 67 Chinese instruments) were identified. Forty-three English instruments were included. Five focused on knowledge and literacy, 12 on belief and self-management perception-related constructs, 21 on self-management and behaviors, and 5 on composite measures. We further recommended 19 English scales with relatively good quality and are frequently applied. Twenty-five Chinese instruments were included, but none were recommended because of a lack of sufficient psychometric property data. CONCLUSION:Many English instruments measuring diabetes self-management have been developed and validated. Further research is warranted to validate instruments adapted or developed in the Chinese population.
With an aging population and a gap in demand for professional caregivers, China's elderly care system is facing severe challenges. Using data mining to expand the smart senior care scenario can effectively improve elderly care services. Based on smart mattress datasets, this paper uses machine learning classification and unsupervised anomaly detection models to analyze the possible risks in the daily behavior of older people from the perspectives of both sleep apnea problems and abnormal physiological information. The model results show that the Stacking algorithm based on data fusion can effectively identify the risk of sleep apnea. In contrast, the Prophet and DBSCAN models can carry out anomaly mining of physiological information of single and combined variables, respectively. Ultimately, based on the research, this paper provides targeted recommendations regarding data collection and integration applications.
A corporate's reputation may be significantly impacted by how it handles emergencies. To investigate the effect of different coping strategies on corporate image restoration, we collected 47 representative events of corporates in different industries as well as relevant comments from Weibo. Based on the image restoration theory proposed by Benoit, the effects of different strategies on corporate reputation restoration were analyzed by using sentiment analysis and the topic model. We found that the two strategies of corrective action and mortification can usually better alleviate the negative emotions of the public. On the contrary, if corporates respond with strategies of denial, evading responsibility, or reducing offensiveness, there could be strong feelings of dissatisfaction. The case study of the food industry also finds that even if the effect of alleviating the negative emotions of the public is not significant, an appropriate response could shift the attention of some consumers and to some extent serve the purpose of restoring the corporate reputation.
The DBSCAN algorithm is a well-known cluster method that is density-based and has the advantage of finding clusters of different shapes, but it also has certain shortcomings, one of which is that it cannot determine the two important parameters Eps ( neighborhood of a point) and Mints (minimum number of points) by itself, and the other is that it takes a long time to traverse all points when dataset is large. In this paper, we propose an improved method which is named as K-DBSCAN to improve the running efficiency based on self-adaptive determination of parameters and this method changes the way of traversing and only deals with core points. Experiments show that it outperforms DBSCAN algorithms in terms of running time efficiency.
[目的/意义]本体广泛应用于知识图谱、信息检索等众多领域,逐渐成为近年来一大研究热点.文章归纳总结了国内外本体构建及应用方面研究成果,以期为相关领域学者提供研究参考.[方法/过程]首先,从本体构建原则、表示语言和构建方法三个维度对本体构建基础研究进行梳理.其次,运用关键节点识别和社团挖掘算法对本体构建主题进行演化分析,并进一步结合中国知网和Web of Science代表性文献梳理近年来本体构建研究热点.最后,结合前述分析总结本体研究主要应用领域,并对本体相关研究进行总结与展望.[结果/结论]结合本体研究现状,半自动化及自动化本体构建、本体融合等技术研究以及本体的应用有望成为未来研究热点.
The stock market, one of the most important trading platforms, has been subject to many market manipulation practices since its emergence. At present, the phenomenon of market manipulation has become an obstacle that restricts the development of the stock market. Based on the data fusion background, this paper conducts a study on the A-share listed companies publicly investigated and confirmed as market manipulation by CSRC in China from 2018-2022 as a manipulation sample under the perspective of fusion of data, model and decision, to provide theoretical support for regulators to identify market manipulation under the new trend. The results show that: ① Comparing with single financial data, different types of models respond to text indicators significantly differently. More than that, a reasonable addition of text dimensions can optimize the model performance; ② Comparing with a single base classifier, integrated models generally have better recognition performance. In addition, this effect can be deeply improved with the diversification of the base classifier.