The role of gut microbiome in predicting diet response and developing personalized dietary recommendations has been increasingly recognized. Yet, we still lack comprehensive, genome-based insights into which gut microbes metabolize specific dietary compounds. Here, we leveraged the metabolic networks constructed from well-annotated microbial genomes to characterize the potential interactions between microbes and metabolites, specifically emphasizing the interactions between microbes and dietary compounds. We revealed a substantial, approximately fourfold variation in both the number of metabolites and dietary compounds in the microbial genome-scale metabolic networks across different genera, whereas species within the same genus showed a high metabolic similarity (mean coefficient of variation in microbial network degree CV = 0.023 for metabolites and 0.015 for dietary compounds). We found that the number of species that can utilize a metabolite drastically varies, ranging from 1 to 818 species, with some metabolites being used by a wide range of species (211 out of 1390 metabolites used by more than 95
Background:Reliable nutrient profiling and semantic interoperability are essential for scalable dietary assessment, food labeling (e.g., traffic-light schemes), and FAIR integration of food composition and consumption data. However, general-purpose large language models (LLMs) are not systematically exposed to structured recipe-nutrition mappings and food ontologies, limiting their accuracy and trustworthiness in food and nutrition tasks. Scope and approach:We review recent LLM advances in life sciences and healthcare and analyze the gap in food and nutrition applications. To address this gap, we introduce FoodyLLM, a domain-specialized LLM fine-tuned on 225k task-aligned QA pairs for (i) recipe nutrient estimation, (ii) traffic-light classification, and (iii) ontology-based entity linking to support FAIR food data interoperability. We benchmark FoodyLLM against strong general-purpose baselines (e.g., Llama 3 8B, Gemini 2.0) under zero-/few-shot prompting across five evaluation folds. Key findings:Across all tasks, FoodyLLM substantially outperforms general-purpose LLMs for nutrient estimation across all macronutrients (fat, protein, salt, saturates, sugar), accuracy increases from 0.43 to 0.63 to 0.91-0.97; for traffic-light classification across all nutrients and color categories, macro F1 improves from 0.46 to 0.80 to 0.86-0.97; and for ontology-based food entity linking across FoodOn, SNOMED-CT, and Hansard, macro F1 increases from 0.33 to 0.44 (best general-purpose baseline) to 0.93-0.98 on artificial NEL data, and from 0.24 to 0.51 to 0.67-0.84 on real corpora (CafeteriaSA and CafeteriaFCD). Overall, our results demonstrate the practical value of domain-specialized LLMs in food and nutrition research. They enable automated dietary assessment, large-scale nutritional monitoring, and FAIR data integration, while opening new pathways toward sustainable and personalized nutrition.
The offering of grocery stores is a strong driver of consumer decisions. While highly processed foods such as packaged products, processed meat and sweetened soft drinks have been increasingly associated with unhealthy diets, information on the degree of processing characterizing an item in a store is not straightforward to obtain, limiting the ability of individuals to make informed choices. GroceryDB, a database with over 50,000 food items sold by Walmart, Target and Whole Foods, shows the degree of processing of food items and potential alternatives in the surrounding food environment. The extensive data gathered on ingredient lists and nutrition facts enables a large-scale analysis of ingredient patterns and degrees of processing, categorized by store, food category and price range. Furthermore, it allows the quantification of the individual contribution of over 1,000 ingredients to ultra-processing. GroceryDB makes this information accessible, guiding consumers toward less processed food choices. Information on the degree of processing of food items is key for better consumer choices. GroceryDB is a dataset with more than 50,000 food items sold at major grocery stores in the United States that uses big data to provide information on processing—categorized by store, food category and price range.
Background: Artificial intelligence (AI) has shown transformative potential across many scientific fields, including food science. Applications span nutrition, safety, flavor, and sustainability. However, current AI implementations in food science often lack integration with domain expertise, face reproducibility challenges, and are hindered by fragmented datasets and limited benchmarking. Scope and approach: This perspective outlines key challenges and proposes five strategic initiatives to guide the effective and responsible integration of AI in food science. These include embedding domain knowledge into models, establishing transparent and reproducible workflows, adopting benchmarking practices, promoting practical validation, and developing robust data standards and infrastructure. Key findings and conclusions: To fully unlock AI's potential in food science, future research must prioritize domain-aware model development, open science practices, and practical validation. These efforts are critical to enabling reliable, generalizable, and impactful AI tools that address real-world challenges in the food systems.
SUMMARY:Network medicine leverages the quantification of information flow within sub-cellular networks to elucidate disease etiology and comorbidity, as well as to predict drug efficacy and identify potential therapeutic targets. However, current Network Medicine toolsets often lack computationally efficient data processing pipelines that support diverse scoring functions, network distance metrics, and null models. These limitations hamper their application in large-scale molecular screening, hypothesis testing, and ensemble modeling. To address these challenges, we introduce NetMedPy, a highly efficient and versatile computational package designed for comprehensive Network Medicine analyses. AVAILABILITY AND IMPLEMENTATION:NetMedPy is an open-source Python package under an MIT license. Source code, documentation, and installation instructions can be downloaded from https://github.com/menicgiulia/NetMedPy and https://pypi.org/project/NetMedPy. The package can run on any standard desktop computer or computing cluster.
The offering of grocery stores is a strong driver of consumer decisions, shaping their diet and long-term health. While highly processed food like packaged products, processed meat, and sweetened soft drinks have been increasingly associated with unhealthy diet, information on the degree of processing characterizing an item in a store is not straight forward to obtain, limiting the ability of individuals to make informed choices. Here we introduce GroceryDB, a database with over 50,000 food items sold by Walmart, Target, and Wholefoods, unveiling how big data can be harnessed to empower consumers and policymakers with systematic access to the degree of processing of the foods they select, and the potential alternatives in the surrounding food environment. The extensive data gathered on ingredient lists and nutrition facts enables a large-scale analysis of ingredient patterns and degrees of processing, categorized by store, food category, and price range. Our findings reveal that the degree of food processing varies significantly across different food categories and grocery stores. Furthermore, this data allows us to quantify the individual contribution of over 1,000 ingredients to ultra-processing. GroceryDB and the associated http://TrueFood.Tech/ website make this information accessible, guiding consumers toward less processed food choices while assisting policymakers in reforming the food supply.
Background/Objectives/Methods: Current research on the link between diet and stroke or myocardial infarction primarily focuses on individual food items. However, people’s eating habits involve complex combinations of various foods. By employing an innovative approach known as the Gaussian graphical model to identify dietary patterns along with the Cox proportional model, the study aimed to identify dietary networks and explore their relationship with the incidence of stroke and/or myocardial infarction in the Korean population. The research utilized data from 84,729 participants in the Korean Genome and Epidemiological Study (KoGES), including the HEXA cohort (61,140 participants), CAVAS cohort (15,419 participants), and Ansan-Ansung cohort (8170 participants). Results: The network identified five dietary patterns or communities consisting of different food groups, while nine food groups did not belong to any community. The High-Protein and Green Tea Community consistently reduced the risk of stroke and myocardial infarction (MI), particularly among females. In most communities, no significant associations with stroke risk were noted in males, and the Rice and High-Calorie Beverages Community was linked to an increased risk of MI in both the total population and females. Conclusions: Dietary patterns derived from network analysis revealed distinct dietary habits in the Korean population, offering new insights into the relationship between diet and the risk of stroke and MI.
This chapter explores the evolution, classification, and health implications of food processing, while emphasizing the transformative role of machine learning, artificial intelligence (AI), and data science in advancing food informatics. It begins with a historical overview and a critical review of traditional classification frameworks such as NOVA, Nutri-Score, and SIGA, highlighting their strengths and limitations, particularly the subjectivity and reproducibility challenges that hinder epidemiological research and public policy. To address these issues, the chapter presents novel computational approaches, including FoodProX, a random forest model trained on nutrient composition data to infer processing levels and generate a continuous FPro score. It also explores how large language models like BERT and BioBERT can semantically embed food descriptions and ingredient lists for predictive tasks, even in the presence of missing data. A key contribution of the chapter is a novel case study using the Open Food Facts database, showcasing how multimodal AI models can integrate structured and unstructured data to classify foods at scale, offering a new paradigm for food processing assessment in public health and research.
The brain has long been conceptualized as a network of neurons connected by synapses. However, attempts to describe the connectome using established network science models have yielded conflicting outcomes, leaving the architecture of neural networks unresolved. Here, by performing a comparative analysis of eight experimentally mapped connectomes, we find that their degree distributions cannot be captured by the well-established random or scale-free models. Instead, the node degrees and strengths are well approximated by lognormal distributions, although these lack a mechanistic explanation in the context of the brain. By acknowledging the physical network nature of the brain, we show that neuron size is governed by a multiplicative process, which allows us to analytically derive the lognormal nature of the neuron length distribution. Our framework not only predicts the degree and strength distributions across each of the eight connectomes, but also yields a series of novel and empirically falsifiable relationships between different neuron characteristics. The resulting multiplicative network represents a novel architecture for network science, whose distinctive quantitative features bridge critical gaps between neural structure and function, with implications for brain dynamics, robustness, and synchronization. ### Competing Interest Statement A.-L.B. is the scientific founder of Scipher Medicine, Inc., which applies network medicine to biomarker development.
Over the past two decades, network medicine (NM) has evolved to help define disease mechanisms, identify drug targets, and guide increasingly precise therapies. In recent years, the integration of NM with artificial intelligence (AI), particularly deep learning techniques, has evolved with increasing applications. AI techniques help elucidate complex disease mechanisms and define precise therapies. The depth of useful, mechanistic information implicit in molecular interaction networks and prior deep learning successes provide a rational basis for combining NM and AI in the analyses of large multiomic datasets to enhance the speed, predictive precision, and biological insights of the computational process. In this review, we provide a summary of concepts related to the combined use of AI and NM as a path to precision medicine, illustrating the success of this joint approach to biomedical complexity and its ongoing challenges.
Nutrition is a significant factor in determining our health that is directly under our control, affecting our risk of chronic conditions like diabetes, heart disease, and cardiovascular diseases. Yet, the nutritional recommendations are centered around 150 essential micro- and macro-nutrients involved in generating energy, forming the basis of our current knowledge of how food affects health. This narrow focus means that the vast majority of food compounds remain unknown and untracked, called the ``dark matter of nutrition.'' Thus, the ability to understand how foods modulate our health is limited, providing little insight beyond the essential nutrients. Here, we review the efforts to map the biochemical composition of food and unveil their impacts on human health. We discuss the current resolution of food composition and the potential of mass spectrometry experiments to improve our knowledge of food compounds. By using a network medicine framework, we show that the possible health associations of food biochemicals can be predicted. Finally, we discuss the potential importance of using machine learning and artificial intelligence techniques in both identifying compounds within food and identifying potential health implications.
Diet, a modifiable risk factor, plays a pivotal role in most diseases, from cardiovascular disease to type 2 diabetes mellitus, cancer, and obesity. However, our understanding of the mechanistic role of the chemical compounds found in food remains incomplete. In this review, we explore the "dark matter" of nutrition, going beyond the macro- and micronutrients documented by national databases to unveil the exceptional chemical diversity of food composition. We also discuss the need to explore the impact of each compound in the presence of associated chemicals and relevant food sources and describe the tools that will allow us to do so. Finally, we discuss the role of network medicine in understanding the mechanism of action of each food molecule. Overall, we illustrate the important role of network science and artificial intelligence in our ability to reveal nutrition's multifaceted role in health and disease.
The binding interactions between small molecules and proteins are the basis of cellular functions. Yet, experimental data available regarding compound-protein interaction is not harmonized into a single entity but rather scattered across multiple institutions, each maintaining databases with different formats. Extracting information from these multiple sources remains challenging due to data heterogeneity. Here, we present CPIExtract (Compound-Protein Interaction Extract), a tool to interactively extract experimental binding interaction data from multiple databases, perform filtering, and harmonize the resulting information, thus providing a gain of compound-protein interaction data. When compared to a single source, DrugBank, we show that it can collect more than 10 times the amount of annotations. The end-user can apply custom filtering to the aggregated output data and save it in any generic tabular file suitable for further downstream tasks such as network medicine analyses for drug repurposing and cross-validation of deep learning models.
To identify healthy, impactful, and equitable foods, we combined health scores from six diverse nutrient profiling systems (NPS) into a meta-framework (meta-NPS) and paired this with dietary guideline adherence assessment via multilevel regression and poststratification. In a case-study format, a commonly debated beverage formulation - 100% orange juice (OJ) - was chosen to showcase the utility and depth of our framework, systematically scoring high across multiple food systems (i.e. a Meta-Score percentile = 93rd and Stability percentile = 75th) and leading to an expected increase of US dietary fruit guideline adherence by & SIM;10%. Moreover, the increased adherence varies across the 300 sociodemographic strata, with the benefit patterns being sensitive to absolute or relative quantification of the difference of adherence affected by OJ. In sum, the adaptable, integrative framework we established deepens the science of nutrient profiling and dietary guideline adherence assessment while shedding light on the nuances of defining equitable health effects.
Studying human dietary intake may help us identify effective measures to treat or prevent many chronic diseases whose natural histories are influenced by nutritional factors. Here, by examining five cohorts with dietary intake data collected on different time scales, we show that the food intake profile varies substantially across individuals and over time, while the nutritional intake profile appears fairly stable. We refer to this phenomenon as 'nutritional redundancy' and attribute it to the nested structure of the food-nutrient network. This network enables us to quantify the level of nutritional redundancy for each diet assessment of any individual. Interestingly, this nutritional redundancy measure does not strongly correlate with any classical healthy diet scores, but its performance in predicting healthy aging shows comparable strength. Moreover, after adjusting for age, we find that a high nutritional redundancy is associated with lower risks of cardiovascular disease and type 2 diabetes.
Link prediction is a core task in graph machine learning, as it is useful in many application domains from social networks to biological networks. Link prediction can be performed under different experimental settings: (1) transductive, (2) semi-inductive, and (3) inductive. The most common setting is the transductive one, where the task is to predict whether two observed nodes have a link. In the semi-inductive setting, the task is to predict whether an observed node has a link to a newly observed node, which was unseen during training. For example, cold start in recommendation systems requires suggesting a known product to a new user. We study the inductive setting, where the task is to predict whether two newly observed nodes have a link. The inductive setting occurs in many real-world applications such as predicting interactions between two poorly investigated chemical structures or identifying collaboration possibilities between two new authors. In this paper, we demonstrate that current state-of-the-are techniques perform poorly under the inductive setting, i.e., when generalizing to new nodes, due to the overlapping information between the graph topology and the node attributes. To address this issue and improve the robustness of link prediction models in an inductive setting, we propose new methods for designing inductive tests on any graph dataset, accompanied by unsupervised pre-training of the node attributes. Our experiments show that the inductive test performances of the state-of-the-art link prediction models are substantially lower compared to the transductive scenario. These performances are comparable, and often lower than that of a simple multilayer perceptron on the node attributes. Unsupervised pre-training of the node attributes improves the inductive performance, hence the generalizability of the link prediction models.
Link prediction is a crucial task in graph machine learning with diverse applications. We explore the interplay between node attributes and graph topology and demonstrate that incorporating pre-trained node attributes improves the generalization power of link prediction models. Our proposed method, UPNA (Unsupervised Pre-training of Node Attributes), solves the inductive link prediction problem by learning a function that takes a pair of node attributes and predicts the probability of an edge, as opposed to Graph Neural Networks (GNN), which can be prone to topological shortcuts in graphs with power-law degree distribution. In this manner, UPNA learns a significant part of the latent graph generation mechanism since the learned function can be used to add incoming nodes to a growing graph. By leveraging pre-trained node attributes, we overcome observational bias and make meaningful predictions about unobserved nodes, surpassing state-of-the-art performance (3X to 34X improvement on benchmark datasets). UPNA can be applied to various pairwise learning tasks and integrated with existing link prediction models to enhance their generalizability and bolster graph generative models.
Piero Fariselli合作论文数University of Bologna2