In Machine Learning (ML) the analysis and the preparation of data before their use is considered an important task, in order to improve the performance of the ML algorithms. Techniques like clustering and dimensionality reduction are applied for the preparation of the data offering several advantages for the ML algorithms to which the data will be used. Some indicative advantages include the improved performance and the faster training of the ML algorithms, the handling of missing or corrupted data, the detection of data overfitting and the visualization of the data. In this paper, a dataset that contains plenty of ship engine's data is analyzed. Specifically, methodologies for density estimation, clustering and dimensionality reduction are studied. Subsequently, such methodologies are applied to the aforementioned dataset, providing useful results about the structure of the dataset, about the correlation of its data, as well as about the importance of each feature included in the dataset.
更多
查看译文
关键词
Machine Learning (ML),Ship Engine Data Analysis,Density Estimation,Data Clustering,Dimensionality Reduction