
Oral historical archive resources are an emerging archive resource with the rapid development of modern technology. Its "bottom-up" approach to historical research has received widespread attention in the fields of history, archives, and libraries. Under the common knowledge discovery mode, oral historical archives resources are showing a dispersed state. Information technology represented by knowledge graphs can break through the data solidification of oral historical archives, reshape the information stack of oral historical archives, and achieve knowledge association and aggregation of oral historical archive resources. The article attempts to construct a knowledge graph of the oral historical archives resources on the theme of "science and art" in the collection of T.D. Lee Library of Shanghai Jiao Tong University. It uses Large Language Model - Retrieval Augmented Generation (LLM-RAG) for knowledge extraction, and then uses a semantic model for knowledge organization and management. The article attempts to empower humanities with technology, exploring the possibility of combining "digital technology" and "humanities research", extending traditional humanities research methods, breaking down barriers between technology and humanities resources, and providing a new path reference for revealing resource content characteristics, semantic deep correlation, and multi-dimensional knowledge discovery.
This paper proposes implementing online, multi-disciplinary tabletop exercises as an experiential learning approach to better prepare future IT service leaders in cybersecurity policy and governance. The interactive simulations will present scenarios spanning regulatory compliance, incident response, risk mitigation, and strategic planning across various industries. Participants will assume diverse roles (e.g. CISOs, legal counsel, operations managers) and apply critical thinking, problem-solving, and communication skills to address the evolving situations. This innovative pedagogy leverages cloud-based tools to facilitate remote, team-based learning experiences that bridge the gap between theoretical concepts and pragmatic application while fostering crucial virtual coordination abilities. The paper outlines the design principles, learning objectives, and anticipated outcomes of these cyber tabletop exercises, aiming to cultivate a holistic cybersecurity mindset that empowers future leaders to develop robust governance frameworks aligning people, processes, and technologies.
Collaborative filtering faces sparsity and cold start problems, while content-based approaches suffer from overspecialization and lack of personalization. Systems such as HSPRec18, HSPRec19, and SemRec have not explored using item images to enrich recommendations. Cross-regional e-commerce systems like Amazon and eBay have retailers from disparate backgrounds. For instance, the same items are likely to be named or described differently by different retailers on these systems, such as “Slippers”, “Flipflop”, and “Sandals” or “Pad” as a female hygiene product or “Pad” as a computer mouse pad. Additionally, systems such as pRNN and Caser attempt to incorporate item images into sequential recommendation using neural network models but suffer from (i) the assumption that adjacent interactions in a sequence must be dependent, which is not always true, (ii) difficulty in training and finding optimal hyper-parameters, (iii) vanishing and exploding gradient problems, and (iv) low interoperability. To improve recommendation accuracy, this paper proposes iHSPRec (Image Enhanced Historical Sequential Pattern Recommendation), which uses items’ image similarity scores to build a vectorized item-item similarity matrix. Then, integrate this with a vectorized sequential pattern of users’ purchases or clicks as input data for the recommendation process. This is achieved by learning the similarity between item image vectors using a convolution neural network, mining visually similar sequential purchase patterns, and enriching the item-item matrix with the products’ structural similarity score and products’ purchase patterns. iHSPRec provides Top-K personalized recommendations based on image similarities between items without needing a user’s rating. Experimental results and comparison with existing systems show that iHSPRec has higher recommendation accuracy than benchmark methods.
Applications of machine learning algorithms encounter Out-of-Distribution (OOD) test cases in the real world, and detection of such samples is a critical problem. The test cases of OOD detection can include near- and far-OODs, making it difficult to address with a single model. We found that OOD detection models, which train exclusively on in-distribution (ID) data, achieve a high performance on far OOD benchmarks, while models trained with augmented or generated data to acquire enriched representation, which is intuitively beneficial for near-OOD benchmarks, can have lower metrics on far OOD benchmarks. In this paper, we propose a norm-based scoring function and a contrastive representation learning to improve near-OOD detection. Furthermore. We propose an ensemble score to take advantage of the proposed model and an ID trained models with superior far OOD detection performance. Our empirical study using a collection of image benchmarks shows the advantage of the proposed ensemble score over the state-of-the-arts for both near- and far-OOD detection benchmarks.
Nudity is also an element in pornography. There are individuals that are conservative about what they are watching. There are children that are too exposed in nudity, and it becomes a normal scene, and it leads to changes of their attitudes. That is why this study developed a system that can detect frontal view of female breast, frontal view of female genital and frontal view of male genital in videos. The researchers tested the accuracy of the system in detecting nude objects in different resolutions. The researchers used experiment method for this research and used formula of Precision, Recall and F-Measure to obtain the accuracy of the system in detecting frontal view of female breast, frontal view of female genital and frontal view of male breast in different resolutions. The system's accuracy got different average precision, average recall and average f-measure in each model depending on their resolution. The higher the resolution is, the higher the accuracy of the model it gets. The researchers recommend the use of relative position of the human body in detecting the breast, female genitals, and male genitals improve the accuracy of the system. For example, the detected breast and genital candidate outside of Human body ROI will be marked as negative. Finally, use other object detection algorithms that can recognize different angles to improve robustness, and other classification algorithms such as Support Vector Machine and Artificial Neural Network.
This work is intended to address the relationship between behavioral economics and decision-making in the within online electronics retail contexts. The work examines the influences of key behavioral economics principles such as Loss Aversion, Social Proof, Endowment Effect, and Status Quo Bias on consumer behavior using a quantitative approach which is applied on a Transactional Retail Dataset. Issues studied include cognitive biases and social factors that shape consumers perceptions of product value, pricing, and satisfaction. Understanding the consumer decision-making in the digital age provides valuable feedback for businesses to help them in the decision making process, and hence they can optimize their marketing strategies, pricing models, and customer experiences. This paper provides a model that can help online electronics retail businesses enhance their effectiveness by understanding behavioral economics parameters.
Digital twins provide conditions for cyber-physical integration, as a bridge connecting the physical world and the cyber world, providing a new way for the manufacturing industry to conduct smart production and precise management. Data from the real world is transmitted to the virtual model through sensors to complete simulation, verification, dynamic adjustment and feedback. By improving designers’ ability to extract knowledge from large and complex data, digital twins can speed up design and development. Therefore, it is necessary to have a deep understanding of how to provide data to designers and how they interact with data. The quality of digital twins would determine their usability and efficiency. This study emphasized on the benefits that digital twins could provide to designers in data processing and integration for the realization of design innovation through the interactive design model of virtual reality. The results of this study could help enhance designers’ ability to control data and strengthen the way designers interact with data. Based on the integrated data environment, designers could conduct interactive analysis, respond to changes in the physical world, improve the design process and add product values.
Vehicle fault prediction is becoming one of the main goals in manufacturers’ maintenance strategies to reduce the number and severity of quality problems in vehicles. Hundreds of vehicle sensors can be used for the early detection of component breakdowns. This work introduces a breakdown prediction approach based on vehicle usage over time. This study proposes a steered optimization system using an evolutionary algorithm called Genetic Algorithm coupled with an Elastic technique to select the most informative predictors. Then, a specific kind of ensemble technique, namely stacking, is utilized for the final prediction. The proposed system has been applied to a complex problem of predictive maintenance to forecast components’ failures. The experimental evaluations on the real usage data collected from thousands of heavy-duty trucks justify the proposed approach is promising.
Cyberbullying poses significant risks in online social media environments, leading to severe psychological distress and societal harm. In this paper, we present a systematic approach to cyberbullying detection, aiming to mitigate these risks and protect individuals in digital spaces. Our approach leverages curated datasets from social media platforms and employs preprocessing techniques to extract vital information and calculate confidence scores for indicators of bullying and aggression. We then utilize an open-source Large Language Model (LLM) to generate specialized detection queries based on the extracted data. Responses generated by the LLM are evaluated for signs of cyberbullying using automated classification methods. Specifically, the presence or absence of key indicators within responses determines their classification. Our systematic approach enables the automated identification and categorization of cyberbullying instances, facilitating proactive intervention and prevention strategies. Through this paper, we contribute to the ongoing efforts to combat cyberbullying and promote safer online interactions.
According to the World Health Organization, Dementia, a chronic, degenerative condition that affects 55 million people worldwide, is prevalent in seniors. To date, there is no cure for dementia, so recent advancements in this area emphasize the urgent need for detecting early symptoms to facilitate timely intervention strategies. As dementia starts by damaging neurons in parts of the brain responsible for memory, language, and thinking, the analysis of language sample could potentially offer a promising avenue for detecting subtle cognitive shifts that could allow the possible interventions to slow down the disease's progression and improve the life of the individuals. We document this paper to explore the potential of leveraging speech and text data for detecting early-stage dementia by leveraging semantic, syntactic, and acoustic features. This paper surveys natural language processing (NLP) techniques applied to speech, text, audio, and handwritten data in monolingual and multilingual settings, focusing on their potential to aid in the early detection and understanding of dementia. According to our study, most datasets available in this domain are in English, with the support vector machine being the most frequently used classification method. Interest is also growing in using large language models to identify the signs of cognitive decline based on language patterns.
In the dynamic realm of information technology, aligning educational curricula with market demands is crucial for producing employable graduates and supporting industry growth. This research addresses the significant lag in the integration of rapidly evolving technologies, such as artificial intelligence and machine learning, into academic programs and identifies discrepancies in skill requirements across various sectors, including cybersecurity. Utilizing a novel approach, we developed a software tool to extract data from extensive job postings on a leading job advertisement platform, www.highered.com. Our analysis provides a detailed examination of the IT job market, highlighting geographical distribution of opportunities, required technical skills, experience levels, and necessary credentials across different sectors. The findings reveal skill mismatches and offer critical insights for educators to update course offerings and tailor educational programs to better meet industry needs. This paper also serves as a guide for policymakers in resource allocation and infrastructure development, aiming to enhance graduate employability and address sector-specific demands effectively. The study not only highlights the importance of timely curriculum updates but also demonstrates the benefits of leveraging real-time labor market data to inform educational strategies.
With the invention of advanced technologies, there are millions of options to improve the quality of life in an urban city. Several innovative implementations transform urban cities into smart cities using new technologies to enhance urban inhabitants' efficiency, sustainability, and overall quality of life. Our study shows that knowledge graphs play an important role in smart cities for transportation, parking, traffic, and city development. They serve as significant repositories, bringing together data from various sources. Several crucial domains of smart cities use knowledge graphs to resolve challenges that hinder urban development. In this paper, we discuss the applications of knowledge graphs in various smart city areas, identify existing challenges, and propose strategies to enhance the current implementation of knowledge graphs in smart cities. We highlight a few innovations that used knowledge graphs in smart cities, showcasing their versatility. Integrating knowledge graphs into smart cities significantly enhances the efficiency of urban services by consolidating and connecting data from different sources and constructing a graph, aiding in better decision-making.
The field of data science revolves around uncovering patterns, extracting meaningful insights, and facilitating data-driven decision making from existing data sources. This research leverages a unique dataset designed to explore the intricate relationships between stress, empowerment, and various workplace outcomes. By harnessing modern techniques spanning classification, regression, and predictive analysis, the objective is to construct a robust model that can guide management strategies and workplace decisions. Given the distinctive nature of the data, rigorous exploratory data analysis techniques are employed to unveil valuable insights. Subsequently, a comprehensive evaluation of multiple classification models is undertaken, including Nearest Neighbors, Linear Support Vector Machines, Radial Basis Function Support Vector Machines, Gaussian Process, Decision Tree, Random Forest, Multi-Layer Perceptron, AdaBoost, and Naive Bayes. The findings provide invaluable business intelligence, elucidating how factors such as accountability, personality traits, and authority-sharing dynamics are likely to influence individual employees' workplace outcomes. Notably, the results demonstrate that Linear Discriminant Analysis, applied in conjunction with multi-dimensional reduction methods, yielded the most accurate predictions among the evaluated models.
Researchers often encounter significant hurdles when dealing with datasets that contain a vast number of missing values. This predicament forces them to make a tough choice: either discard a substantial amount of data, which could drastically undermine the accuracy of the machine learning (ML) model, or attempt to fill these missing values in sensitive medical datasets—a method that is far from ideal. This paper proposes an approach to this issue, suggesting that bypassing the traditional path of data imputation in favor of a model that learns from the missing values themselves could paradoxically improve the accuracy and predictive capabilities of Alzheimer's Disease (AD) identification models. We introduce a comparison between state-of-the-art ML models and the XGBoost algorithm, which is designed to integrate the learning of missing values into its training cycle, using the official ADNI datasets with extensive missing values. The experiment further evaluates these models on the same datasets post-imputation. The results strikingly indicate that this unconventional strategy not only bridges the gaps created by missing data but also surpasses the accuracy of traditional methods that rely on filling in incomplete samples. This discovery opens up new avenues for research in medical diagnostics for conditions like AD, where data scarcity and imperfections are common. By rethinking how we handle incomplete data, we unlock new potential for refining ML applications in healthcare, particularly in enhancing the precision of diagnoses in complex diseases such as AD.
The verification and validation of software system class models are crucial aspects of software engineering, aiming to ensure correctness, reliability, and quality in software systems. This research explores advancements in verification and validation techniques, focusing on UML class diagram models. The study proposes a methodology to enhance the reliability and security of software models by simplifying UML class diagrams, eliminating semantic overload, and validating models using the Object Constraint Language (OCL). The practical and theoretical values of this research are discussed, highlighting benefits for developers, system administrators, end-users, cybersecurity professionals, organizations, and academic communities. The work addresses challenges and limitations in current verification and validation techniques, emphasizing the complexity of modern software systems. The proposed methodology is implemented using the USE tool, demonstrating its effectiveness in validating UML-based specifications and OCL constraints.
This paper presents a distinctive approach for evaluating the quality of mobile apps in Turkish banks. The assessment is conducted using the Analytic Hierarchy Process (AHP) method, setting it apart from traditional methods. The criteria for app evaluation were determined through interviews with 10 banking industry experts, resulting in a comprehensive list of features considered essential for banking apps in Turkey. Additionally, input from banking managers was gathered to determine the most important criteria and their relative importance. By applying the AHP method, the weight of each criterion was calculated, incorporating both qualitative and quantitative measures. The study stands out for its novel approach to evaluating mobile app quality, specifically in Turkish banks. By utilizing the AHP method and incorporating expert perspectives, the evaluation provides a robust and comprehensive assessment of app features. The findings have the potential to drive improvements in mobile app development and enhance user satisfaction within the Turkish banking industry.
Worldwide sports activities have one of the largest supporter/fan bases of all the many areas of entertainment. The sport of soccer is one of the richest sports in the world and commands the position of being the most popular sport in the world. That single data point forms the basis for an interesting set of data analytical research questions. First on that list would be: How is soccer determined to be the most popular sport in the world? How do the factors of advertising/marketing, televising/streaming, teams’ win/loss, players, location, etc. contribute to the popularity of the sport? What impact does the popularity of soccer have on the players’ performance? Can data analytics predict teams’ performance? The report presented herein is a documentation and review of prior research efforts that have applied soft computing methods to predict teams’ win/loss probability in the sport of soccer. This review comments on the success of the data analytics applied and makes an assessment of how this can affect the popularity of the sport.
This paper presents a two-stage prediction model designed to effectively address the fluctuations in sales volume experienced by supermarket chains due to seasonal patterns, sales cycles, holidays, and irregular variations. In the initial phase, the model combines raw sales data with Fast Fourier Transform (FFT) to uncover seasonal characteristics and then employs a Genetic Algorithm (GA) with its binary encoding capabilities for feature selection. In the subsequent stage, the model leverages the Light Gradient Boosting Machine (LightGBM) as the prediction model, using weak learning strategies to manage irregular variations in sales data. Empirical simulations using three publicly available datasets of supermarket chains from Kaggle validate the effectiveness of our proposed method. This approach holds significant practical implications for understanding and predicting fluctuations in supermarket sales data and offers a new direction for future research in other domains.
Skeleton-based action recognition has garnered widespread attention. However, due to the inherent limitations of skeleton sequences, existing works often confuse actions with inter-class similarities and struggle to meet the requirement for viewpoint invariance. As a solution, multimodal action recognition leverages the complementarity of information between modalities to significantly enhance the performance of unimodal models. However, effectively integrating these modalities remains an open problem. In this work, we first propose a keypoints-based multimodal data fusion method to construct images that adequately represent the crucial spatiotemporal characteristics and their variations of actions. Building upon this, we introduce the keypoints-based multimodal fusion network (KBMN), which comprehensively learns action features from skeleton, RGB, and depth data. Extensive experiments on two large-scale datasets demonstrate that our KBMN exhibits robust performance in both unimodal and multimodal action recognition tasks. As an auxiliary model for skeleton-based methods, KBMN effectively assists various baseline methods in improving their recognition accuracy.
An Opportunistic Network is a revolutionary technology making a huge impact on human society. It gained significant attention from researchers, due to its potential application domains like Vehicular Networks, Delay Tolerant Networks, or Internet of Things, etc. In case of an Opportunistic Network, whenever the source node transfers packets to destination node in the presence of malicious nodes, the malicious nodes drops few legitimate packets and inject malicious(fake) packets to launch attacks. Several techniques have been proposed by researchers to mitigate the issue, however hackers are ways to overcome these new approaches and still find a ways to compromise the communication. Due to this research, this research propose a modified algorithm based on our previously proposed algorithm FAPMIC. We used Cyber-Threat-Intelligence Strategy to take preventive measure against malicious nodes which launch selective packet drops and fake packet attacks simultaneously. The proposed modified algorithm enhances the detection accuracy, detection delay, false positive/false negative rates, and resources consumption and also improves packet delivery ratios/packet loss ratios.