Taste perception is a crucial factor in the development of innovative strategies for pharmaceutical and food design, as represented by the five primary tastes such as sweet, bitter, umami, salty, and sour. However, relying on human sensory evaluation for taste assessment is time-consuming and costly. Consequently, developing the five tastes predictive model presents a more efficient and practical alternative. Key challenges include the multi-label nature of taste perception, whereby a single compound can exhibit more than one basic taste, severe class imbalance, and the integration of structurally distinct small molecules and peptides within a unified predictive framework. To address this issue, a novel taste predictive model was proposed, incorporating molecular structure analysis and natural language processing. Additionally, tree-based machine learning models, including random forest, XGBoost, LightGBM, and CatBoost, were trained based on feature extraction techniques such as descriptors, fingerprints, FastText, and Word2vec. Experimental results indicated that the proposed model significantly improved classification performance, with LightGBM achieving the highest F1-score of 74.91%. This research not only advances in five basic tastes predictions but also provides key insights of crucial factors for understanding molecular structure of taste perception. These findings are intended to support researchers in computational food science and those involved in the development of healthier food products.
更多
查看译文
关键词
Five Basic Tastes Predictions,Group-aware Data Splitting,Molecular Structures,Multi-label Classification,Natural Language Processing