The reliability of a ship's main engine is essential for uninterrupted maritime operations. Yet existing anomaly detection approaches remain limited: most rely solely on internal engine sensor data and thus fail to account for environmental influences, while others mix internal and external factors without distinction, making it difficult to disentangle outcomes from causal drivers. Moreover, the prevalent use of black-box models hinders interpretability and limits practical diagnosis. To overcome these limitations, this study introduces a novel methodology that integrates a Variational Autoencoder (VAE), the Root Cause Anomaly Predictor (RCAP), and Shapley Additive Explanations (SHAP). The VAE detects anomalies from internal engine sensor data, while RCAP, implemented with LightGBM, predicts anomaly scores using only external operational and environmental variables. SHAP is then applied to RCAP to clearly explain how each external factor contributes to anomaly scores, enhancing interpretability. The methodology was validated using 18 months of operational data from a container ship, demonstrating robustness under real-world conditions. From this dataset, 335 anomalies were detected and categorized into four distinct types through hierarchical clustering of internal sensor patterns, and RCAP-SHAP analysis successfully revealed their associations with external factors such as wind wave height, draft trim, and sea surface temperature. By explicitly separating anomaly manifestation from external causal drivers, this integration of VAE with RCAP and SHAP provides a novel and interpretable approach to anomaly detection, offering actionable insights for proactive maintenance, adaptive routing, and operational cost reduction.
Deep learning (DL) applications in seismic data processing is crucial for accurate subsurface interpretation. However, post-migration seismic data often suffer from noise that can obscure key geological features and impair analysis. Despite the availability of various denoising techniques, a significant gap remains in achieving both robust performance and interpretability without reliance on extensive labeled datasets. This paper proposes a novel solution to this emerging issueby employing an explainable self-supervised approach that combines Deep Convolutional Denoising (DCD) networks with a SHAP-based Noise Contamination detection model and isolation forest techniques. Unlike existing models, which are limited by labeled datasets, our proposed approach emphasizes on the noise features while using DCD's predictive capabilities. Our findings highlighted the novel use of the SHAP denoising approach to effectively isolates and modifies seismic noise. The model was validated using a real world seismic dataset and proven to perform exceptionally well when compared to traditional approaches. Using the Peak Signal-to-Noise Ratio (PSNR) as the benchmark metric, the SHAP 15% contamination approach fared better at various noise scales. Notably, with a White Gaussian Noise Scale of 60.0, the PSNR method result was 24.62, above the scores of the Random Noise contamination (24.39) and Ground Roll contamination approaches (23.46). Similar superiorities were observed at scales 40.0 and 80.0, reaffirming the SHAP 15% contamination method's superiority. Consequently, this integrated approach not only promises enhanced seismic denoising but also demonstrates quantifiable and explainable superior performance metrics, laying the foundation for future seismic processing endeavors.
Fuel oil consumption (FOC) in vessels is influenced by various factors, with vessel load conditions being a critical determinant. This research aimed to develop an interpretable regression-based machine learning model, specifically an XGBoost Regressor, to predict FOC, and leveraged the analysis with Explainable Artificial Intelligence (XAI) techniques to enhance transparency and understanding of the factors affecting FOC in maritime operations. Hyperparameter tuning optimized model parameters, achieving a high R-squared (R2) value of 0.99. XAI methods clarified how operational (e.g., speed, load, draft) and environmental factors (e.g., wind, wave, current, sea state) contribute to FOC increases. As a result, operational factors, notably average draft, exerted a substantial influence, with an average draft of 18.425 m resulting in a significant increase in predicted FOC to 1,729.40 kg/h during the laden condition, compared to the base value of 1,496.82 kg/h. Additionally, environmental factors, particularly Relative Wind Angle, significantly impacted FOC prediction, with a Mean Absolute SHAP Value of 11.97 kg/h, notably higher than in ballast (2.59 kg/h) and empty (2.85 kg/h) conditions. These findings emphasize the importance of tailored fuel efficiency strategies, as operational and environmental factors may vary in their impact across different loading conditions. 'This article is a revised and expanded version of a paper entitled "Predictive Analysis of Fuel Oil Consumption in Vessels: Interpretable Modeling with Emphasis on Load Conditions" presented at The 11th International Conference on Logistics and Maritime Systems (LOGMS 2023) on 6 September 2023 at the Busan Port International Exhibition & Convention Center (BPEX)'
International shipping is responsible for approximately 2.7% of the global greenhouse gas emissions, a share expected to rise by as much as 250% by 2050. In response, the International Maritime Organization (IMO) has set ambitious targets to reduce these emissions to near-zero by 2050, focusing on alternative fuels like LNG. This study examines the energy consumption patterns of dual-fuel engines powered by LNG and develops machine learning models using LightGBM to predict fuel usage for both fuel oil (FO) and gas (GAS) modes. The methodology involved analyzing operational data to identify patterns in fuel usage across different voyage conditions. The FO mode was found to be predominantly used for rapid propulsion during speed changes or directional shifts, while the GAS mode was optimized for stable conditions to maximize fuel efficiency. Additionally, a mixed mode of FO and GAS was occasionally applied on complex routes to balance safety and efficiency. Using these insights, LightGBM models were trained to predict fuel consumption in each mode, achieving high accuracy with R2 scores of 0.94 for the GAS mode and 0.98 for the FO mode. This model enables ship operators to optimize fuel decisions in response to varying voyage conditions, resulting in reduced overall fuel consumption and lower CO2 emissions. By applying the predictive model, operators can adjust fuel usage strategies to match operational demands, potentially achieving notable cost savings and meeting stricter environmental regulations. Furthermore, the accurate estimation of fuel usage supports CO2 emissions management, aligning with the Carbon Intensity Indicator (CII) and providing ship operators with actionable data for fleet management optimization. This research provides essential data to support carbon emission compliance, improves fuel efficiency, and offers practical insights into fuel management strategies. The predictive model serves as a valuable resource for ship operators to optimize fuel use and aligns with the IMO’s environmental targets, aiding the maritime industry’s transition toward carbon neutrality.
In this study, we employed ChatGPT, an advanced large language model, to analyze hotel reviews, focusing on aspect-based feedback to understand service failures in the hospitality industry. The shift from traditional feedback analysis methods to natural language processing (NLP) was initially hindered by the complexity and ambiguity of hotel review texts. However, the emergence of ChatGPT marks a significant breakthrough, offering enhanced accuracy and context-aware analysis. This study presents a novel approach to analyzing aspect-based hotel complaint reviews using ChatGPT. Employing a dataset from TripAdvisor, we methodically identified ten hotel attributes, establishing aspect–summarization pairs for each. Customized prompts facilitated ChatGPT’s efficient review summarization, emphasizing explicit keyword extraction for detailed analysis. A qualitative evaluation of ChatGPT’s outputs demonstrates its effectiveness in succinctly capturing crucial information, particularly through the explicitation of key terms relevant to each attribute. This study further delves into topic distributions across various hotel market segments (budget, midrange, and luxury), using explicit keyword analysis for the topic modeling of each hotel attribute. This comprehensive approach using ChatGPT for aspect-based summarization demonstrates a significant advancement in the way hotel reviews can be analyzed, offering deeper insights into customer experiences and perceptions.
First, we propose a class of efficient models classed as choice-based recommendation (CBR) for parametric metrics, such as a logit model as a recommendation system using nonparametric approaches. The rest of the papers is organized as follow : we used a simple, streamlined architecture that uses a nonparametric approach such as a feedforward deep neural network (DNN). The study implemented a method to deal with a choice set with a fixed and variable-length option, investigate deep learning methods that consider each choice set as one sample point, the effect of embedding categorical features and accuracy impact, and the efficiency of batch normalization toward a more stable network. To check the performance of our approach, we conducted extensive experiments on multiple datasets and used the top-k accuracy as a metric. We then show the effectiveness of CBR across two industrial applications and use cases, including hotel booking and airline itineraries. The results show that the DNN outperforms the multinomial logit model (MNL) with significant top-k accuracy. The top-k accuracy was further divided into three different DNN models. Among the models, a model that included a layer with batch normalization embedding outperforms with top-k accuracy compared with the model that does not include both batch normalization and embedding layer in the proposed DNN architecture.
Well log data imputation is crucial for subsurface geology interpretation, which helps identify the most productive areas for drilling, minimize exploration risks, and maximize hydrocarbon recovery. Obtaining high-quality well data can be challenging due to various issues such as drilling problems, improper logging processes, or operational issues with logging tools. Accurate and reliable imputation methods are essential to address missing or incomplete information in well log data, allowing better decision-making in the oil and gas industry. This study presents a novel well log data imputation method for the West Natuna Basin in Indonesia, using time-series deep learning models. Focusing on the “Kappa” well dataset with missing Vp log data, we trained long short-term memory (LSTM), gated recurrent unit (GRU), and bidirectional LSTM (Bi-LSTM) models using a window-based time step sequence of 10-timesteps intervals as input. The LSTM was observed as the best model, achieving a mean absolute percentage error (MAPE) of about 2.2% and an R-squared (R2) of about 94%. This result suggests that deep sequence model can be effectively used for missing data imputation in well log. Lastly, a petrophysical analysis further underlines the importance of these findings, offering deeper insights into reservoir properties, and value of our research in real-world exploration scenarios. The findings can aid in improving the accuracy and reliability of imputation in well log data, enabling better decision-making in the oil and gas industry.
In the maritime industry, optimizing vessel fuel oil consumption is crucial for improving energy efficiency and reducing shipping emissions. However, effectively utilizing operational data to advance performance monitoring and optimization remains a challenge. An XGBoost Regressor model was developed using a comprehensive dataset, delivering strong predictive performance (R2 = 0.95, MAE = 10.78 kg/h). This predictive model considers operational (controllable) and environmental (uncontrollable) variables, offering insights into complex FOC factors. To enhance interpretability, SHAP analysis is employed, revealing ‘Average Draught (Aft and Fore)’ as the key controllable factor and emphasizing ‘Relative Wind Speed’ as the dominant uncontrollable factor impacting vessel FOC. This research extends to further analysis of the extremely high FOC point, identifying patterns in the Strait of Malacca and the South China Sea. These findings provide region-specific insights, guiding energy efficiency improvement, operational strategy refinement, and sea resistance mitigation. In summary, our study introduces a groundbreaking framework leveraging machine learning and SHAP analysis to advance FOC understanding and enhance maritime decision making, contributing significantly to energy efficiency and operational strategies—a substantial contribution to a responsible shipping performance assessment under tightening regulations.
A vessel sails above the ocean against sea resistance, such as waves, wind, and currents on the ocean surface. Concerning the energy efficiency issue in the marine ecosystem, assigning the right magnitude of shaft power to the propeller system that is needed to move the ship during its operations can be a contributive study. To provide both desired maneuverability and economic factors related to the vessel’s functionality, this research studied the shaft power utilization of a factual vessel operational data of a general cargo ship recorded during 16 months of voyage. A machine learning-based prediction model that is developed using Random Forest Regressor achieved a 0.95 coefficient of determination considering the oceanographic factors and additional maneuver settings from the noon report data as the model’s predictors. To better understand the learning process of the prediction model, this study specifically implemented the SHapley Additive exPlanations (SHAP) method to disclose the contribution of each predictor to the prediction results. The individualized attributions of each important feature affecting the prediction results are presented.
Diabetes was associated with an increased risk of cardiovascular disease in both adult cancer survivors and the general population. The magnitude of association between diabetes and cardiovascular disease was greater in adult cancer survivors than that in the general population. Lay Summary In a matched cohort of adult cancer survivors and the general population, diabetes was associated with a significantly higher risk of cardiovascular disease in adult cancer survivors compared to that in the general population. Aims Diabetes is a well-established risk factor for cardiovascular disease (CVD), but little is known about the differences in contribution of diabetes to incident CVD between adult cancer survivors and those without history of cancer. The aim of this study was to evaluate the magnitude of association between diabetes and CVD risk among adult cancer survivors and their general population counterparts. Methods and results The National Health Insurance Service database was used to abstract data on 5199 adult cancer survivors and their general population controls in a 1:1 age- and sex-matched cohort setting. The Cox proportional hazards model adjusted for socioeconomic status, health status, lifestyle, and clinical characteristics was used to calculate hazard ratios (HR) and 95% confidence intervals (95% CI) of incident CVD associated with glycaemic status in adult cancer survivors and the general population. The partial likelihood ratio test was used to compare the magnitude of the association between diabetes and CVD risk in the two groups. Compared to those without diabetes, adult cancer survivors (adjusted HR = 2.30; 95% CI: 1.24-4.30) and their general population controls (adjusted HR = 1.91; 95% CI: 1.02-3.58) with diabetes had a higher risk of incident cardiovascular outcomes. The magnitude of diabetes-CVD association was significantly stronger in adult cancer survivors than that in those without history of cancer (P = 0.011). Conclusions The magnitude of association between diabetes and incident CVD was stronger in adult cancer survivors as compared to that in their general population counterparts, supporting evidence for the importance of glycaemic control for prevention of CVD among those with history of cancer diagnosis and treatment.
The main engine of a ship plays a crucial role in providing propulsion. In recent times, there has been growing interest in a data-driven monitoring approach that utilizes sensor data to complement the preventive maintenance-centered maintenance strategy. Previous studies have proposed methodologies that apply anomaly detection algorithms to the sensor data within the main engine. However, these methodologies have limitations as they only focus on analyzing internal sensor data and fail to consider external factors such as operating conditions, marine environment, and weather. Additionally, the use of black-box approaches makes it challenging to determine the specific factors causing anomalies. To address these limitations, this study introduces a method that employs Explainable Artificial Intelligence (XAI) techniques to identify the causes of anomalies in ship main engines. The proposed method involves calculating anomaly scores using Variational AutoEncoder on collected sensor data and training a separate model to predict anomaly scores by considering external factors like operating conditions and weather. Furthermore, the SHAP (Shapley Additive Explanations) technique is utilized to quantify the contributions of external factors to the anomaly scores. This enables the analysis of individual data features and facilitates both local and global analysis for identifying the causes of anomalies and diagnosing faults. The proposed methodology was validated through a case study using data collected from a container ship over an 18-month period, demonstrating its effectiveness in identifying the causes of anomalies in the ship's main engine.
As software systems evolve, they become more complex and larger, creating challenges in predicting change propagation while maintaining system stability and functionality. Existing studies have explored extracting co-change patterns from changelog data using data-driven methods such as dependency networks; however, these approaches suffer from scalability issues and limited focus on high-level abstraction (package level). This article addresses these research gaps by proposing a file-level change propagation to vector (FCP2Vec) approach. FCP2Vec is a recommendation system designed to aid developers by suggesting files that may undergo change propagation subsequently, based on the file being presently worked on. We carried out a case study utilizing three publicly available datasets: Vuze, Spring Framework, and Elasticsearch. These datasets, which consist of open-source Java-based software development changelogs, were extracted from version control systems. Our technique learns the historical development sequence of transactional software changelog data using a skip-gram method with negative sampling and unsupervised nearest neighbors. We validate our approach by analyzing historical data from the software development changelog for more than ten years. Using multiple metrics, such as the normalized discounted cumulative gain at K (NDCG@K) and the hit ratio at K (HR@K), we achieved an average HR@K of 0.34 at the file level and an average HR@K of 0.49 at the package level across the three datasets. These results confirm the effectiveness of the FCP2Vec method in predicting the next change propagation from historical changelog data, addressing the identified research gap, and show a 21% better accuracy than in the previous study at the package level.
Well logs are critical dataset to interpret the subsurface geology as they represent the physical characteristics of the logged formations. At certain intervals however, such information may be missing and/or erroneous due to the drilling problems (i.e. much heavier mud weight as opposed to the formation integrity, leading to formation damage), improper logging process or tools operational issues. Such issues are generally recognized by the presence of anomalous data spikes in the case of logging tools issues or erroneously low log reading in the damaged formation. To address this issue, a new solution based on deep learning approach was tested to imputate such missing value for sonic log (Vp) data. The missing well log values were anticipated using data-driven machine learning methods, particularly the GRU, LSTM, and Bi-LSTM with window-based time step sequence inputs. An empirical research technique was utilized to determine the optimum parameters to achieve the best validation score. The relative value of various input parameters were assessed throughout the training phase in order to eliminate insensitive measures and prioritize data with a high association to the target variables. Two key wells (Zeta and Kappa) from the prolific West Natuna Basin were utilized to test and to validate the deep learning performance. This paper briefly discusses the key findings of this study, pitfalls, and performance analysis in details related to the proposed models.
Introduction: Whether the contribution of elevated low-density lipoprotein (LDL) cholesterol to cardiovascular outcomes differs between adult cancer survivors and general population without history of cancer remains uncertain. Hypothesis: Elevated LDL cholesterol (≥130 mg/dL/≥3.36 mmol/L) is associated with incident cardiovascular disease (CVD) and the magnitude of this association is stronger in adult cancer survivors compared to the general population without history of cancer. Methods: Individuals ≥19 years of age were enrolled in the National Health Insurance Service-National Sample Cohort (NHIS-NSC) in 2002 with follow-up through December 30, 2015. In this study, we included adult cancer survivors without history of CVD and statin treatment who survived more than 1 year after the first cancer diagnosis and their age-and-sex matched control with 1:1 ratio. We calculated hazard ratios (HRs) and 95% confidence intervals (95% CIs) from Cox proportional hazards model adjusted for shared CVD risk factors between adult cancer survivors and the general population. Q-statistic was used to compare the difference in the magnitude of elevated LDL cholesterol-CVD associations between the two populations. Results: Between January, 2003 and December, 2006, 5,163 adult cancer survivors and their matched control (N=5,163) were identified in the NHIS-NSC. The adjusted HRs and 95% CIs for incident CVD among adult cancer survivors and the general population with elevated LDL cholesterol, as compared to those without, were 1.87 (95% CIs: 1.22-2.87) and 1.14 (95% CIs: 0.73-1.76), respectively. However, the difference in contribution of elevated LDL cholesterol to incident CVD between the two populations was not statistically significant ( P heterogeneity =0.114). Conclusion: In this population-based cohort, we found an association between elevated LDL cholesterol and incident CVD, specifically in adult cancer survivors. These findings suggest that adult cancer survivors with elevated LDL cholesterol may need additional clinical attention for cardiovascular health.
Introduction: Little is known about the performance benefit of combining these data with publicly available environmental risk factors for CVD prediction models. Hypothesis: We aimed to test whether routinely collected clinical data (age, sex, systolic blood pressure, total cholesterol, cigarette smoking) in combination with data on environmental risk factors could improve the performance of CVD prediction using an explainable machine learning (ML) technique. Methods: We identified individuals without previous history of CVD from the National Health Insurance Service, 2002-2015 in the Republic of Korea, linked to publicly available data on annual average of fine particulate matter (PM 10 ) exposure and urban green space coverage according to the administrative code for residence in each individual. Random forest (RF) model and SHapley Additive exPlanations (SHAP) value were used for prediction for newly diagnosed CVD and explanation for contribution of each component to the model output, respectively. Performance of the prediction models were evaluated using area under the curve (AUC). Results: Among 151,936 individuals included, there were 2,837 subsequent CVD events. The AUCs for the RF model with the clinical data only and the model with environmental risk factors (high annual average PM 10 and low UGS coverage) in addition to the clinical data were 0.731 and 0.733, respectively. The SHAP values showed that adding these environmental risk factors did not have significant impact on the model output (SHAP summary plot). However, the clinical data comprised of traditional CVD risk factors had notable contribution to the performance of prediction model. Conclusions: This study showed that adding publicly available data on environmental risk factors to the routinely collected clinical data had only marginal improvement in the prediction of CVD outcomes.
This study aims to put a supervised learning method for automatically classifying lithofacies in well-logging dataset, where several machine learning algorithms were compared in this study that took place in the Tarakan Basin, Indonesia. The predicted lithofacies in this study including shale, shaly sandstone, sandstone, and coal, where coal is considered as the unique lithofacies in the study area. As training and testing datasets, we used two separate well log datasets from the Tarakan Basin. The first well, named Omnicron, was used to train the model, while the second well, named Kay, was used to test it. Random Forest and Gradient Boosting outperformed the other models in the experiment, with the accuracy of 87.49% and 87.01%, respectively. When it came to classifying coal, however, both approaches had issues. The Pr-Recall curve revealed that the coal score was under average precision in each facies, with values of roughly 0.52 and 0.38, respectively, which explaining why, even with high accuracy, the machine learning algorithm predicted poorly in one lithofacies class. In order to evaluate this coal misclassification, we used rock physics to analyze the machine learning prediction in this report. As result, we found that each facies is well-differentiated by physical properties, and the predicted lithofacies have a distribution that is close to the original facies however, coal may be potentially misclassified as other lithofacies as some of the coals have similar rock physical properties with the surrounding lithology (e.g. coal with a mixture of shale may have similar DT and GR responses). Based on this research, the use of machine learning in the Tarakan Basin effectively provides lithofacies data with a high degree of precision and accuracy in a much shorter time.
In a large ship or vessel, there are a lot of sensors forming a system that is used to indicate the engine status. It is critical for the system to be able to detect any anomaly that may cause engine failures. By detecting the anomaly of the data, maintenance for the sensors can be well-recommended and this also contributes to the reduction of maintenance costs. In this research, a collection of sensor data from vessels was analyzed using an Isolation Forest to detect the anomaly of the data. To reduce the dimensionality of the data, the t-SNE was adopted.
In this study, we proposed a data-driven approach to the condition monitoring of the marine engine. Although several unsupervised methods in the maritime industry have existed, the common limitation was the interpretation of the anomaly; they do not explain why the model classifies specific data instances as an anomaly. This study combines explainable AI techniques with anomaly detection algorithm to overcome the limitation above. As an explainable AI method, this study adopts Shapley Additive exPlanations (SHAP), which is theoretically solid and compatible with any kind of machine learning algorithm. SHAP enables us to measure the marginal contribution of each sensor variable to an anomaly. Thus, one can easily specify which sensor is responsible for the specific anomaly. To illustrate our framework, the actual sensor stream obtained from the cargo vessel collected over 10 months was analyzed. In this analysis, we performed hierarchical clustering analysis with transformed SHAP values to interpret and group common anomaly patterns. We showed that anomaly interpretation and segmentation using SHAP value provides more useful interpretation compared to the case without using SHAP value.
Introduction: The 2017 American College of Cardiology/American Heart Association (ACC/AHA) guideline changed the definition of hypertension, which has generated newly defined population with hypertension in the United States and elsewhere. However, there is limited evidence for the difference in contribution of hypertension according to the 2017 ACC/AHA guideline to cardiovascular outcomes in adult cancer survivors and general population without history of cancer. Hypothesis: We hypothesized that both adult cancer survivors and the general population with hypertension according to 2017 ACC/AHA definition (blood pressure ≥130/80 mm Hg) would be at higher risk of cardiovascular outcomes compared to those without hypertension and magnitude of this association would be different between the two populations. Methods: We used data from the National Health Insurance Service-National Sample Cohort (2002-2015) to identify adult (≥19 years) cancer survivors who survived more than 1 year after the first-ever cancer diagnosis and those without history of cancer matched for age and sex with 1:1 ratio. In both populations, those with history of cardiovascular disease were excluded. Cox proportional hazards model adjusted for the shared cardiovascular risk factors in the two populations and Q-statistic were used for analyses. Results: In this age and sex matched cohort, we identified 5,163 adult cancer survivors matched to their controls. The adjusted hazard ratio (HR) and 95% confidence intervals (95% CI) for cardiovascular outcomes in adult cancer survivors with hypertension compared to those without hypertension was 1.95 (95% CI: 1.21-3.15). Hypertension was also associated with an increased risk of cardiovascular outcomes in the general population (adjusted HR=1.67, 95%: CI 1.07-2.66). However, magnitude of this association did not differ significantly between the two populations ( P heterogeneity =0.646). Conclusions: Both adult cancer survivors and the general population with hypertension according to the 2017 ACC/AHA guideline are at higher risk of cardiovascular outcomes compared to those without hypertension. Difference in the magnitude of increased risks for cardiovascular outcomes between the two populations was not discernable.