Synthetic tabular data generation has attracted growing attention due to its importance for data augmentation, foundation models, and privacy. However, real-world tabular datasets increasingly contain free-form text fields (e.g., reviews or clinical notes) alongside structured numerical and categorical attributes. Generating such heterogeneous tables with joint modeling of different modalities remains challenging. Existing approaches broadly fall into two categories: diffusion-based methods and LLM-based methods. Diffusion models can capture complex dependencies over numerical and categorical features in continuous or discrete spaces, but extending them to open-ended text is nontrivial and often leads to degraded text quality. In contrast, LLM-based generators naturally produce fluent text, yet their discrete tokenization can distort precise or wide-range numerical values, hindering accurate modeling of both numbers and language. In this work, we propose TabDLM, a unified framework for free-form tabular data generation via a joint numerical–language diffusion model built on masked diffusion language models (MDLMs). TabDLM models textual and categorical features through masked diffusion, while modeling numerical features with a continuous diffusion process through learned specialized numeric tokens embedding; bidirectional attention then captures cross-modality interactions within a single model. Extensive experiments on diverse benchmarks demonstrate the effectiveness of TabDLM compared to strong diffusion- and LLM-based baselines.
In recent years, the rapid advancement of high-throughput technologies has led to the generation of vast and complex multi-omics datasets that are valuable for characterizing and understanding complex cell signaling network systems. On the other hand, large language models (LLMs), domain-specific foundation models (FMs) and AI agents, have achieved significant breakthroughs and have been revolutionizing scientific research. The convergence of these two trends is catalyzing a new era for biomedical research to augment and speed up scientific discovery and the development of precision medicine. In this study, we examine the large-scale omics datasets, emerging applications, and challenges at the intersection of massive omic datasets and related AI models and agents, highlighting how their integration is reshaping the landscape of biomedical research and precision medicine.
Hawthorn wine has gained increasing popularity in China, but comprehensive research on its sensory and chemical characteristics is still limited. This study established a sensory lexicon using Pivot Profile to describe and differentiate Chinese hawthorn wines. Based on the sensory data, 13 hawthorn wines presented three different styles, namely ‘Sweet’, ‘Fruity’ and ‘Alcohol’. A total of 129 volatile compounds were identified and quantified in all hawthorn wines using headspace solid-phase microextraction (HS-SPME) combined with both gas chromatography-Orbitrap mass spectrometry (GC-Orbitrap-MS) and gas chromatography-Quadrupole mass spectrometry (GC-Quadrupole-MS). Partial least-squares regression revealed that ‘sweet’, ‘hawthorn’, and ‘honey’ attributes were positively correlated with several terpenes, volatile phenols, lactones, ethyl cinnamate, nonanal and phenylacetaldehyde, as well as sugar content, while negatively correlated with alcohol content. Furthermore, a salting-out effect of certain terpenes and volatile phenols was observed with increasing sucrose concentration, potentially enhancing the perceived intensity of these above attributes.
Epicotyl dormancy represents a distinct form of seed dormancy in temperate tree species. This physiological mechanism delays shoot emergence relative to root growth until the seeds fulfill their specific heat requirement. While previous research has predominantly focused on seed germination, seedling emergence plays a more critical role in forest ecosystems by directly influencing population establishment. Cold stratification is known to facilitate the release of epicotyl dormancy in some species, yet its impact on the heat requirements for seedling emergence remains poorly understood. In this study, we proposed the heat requirement for the first time in epicotyl dormancy release and emergence and investigated the effects of cold stratification on the heat requirements for epicotyl dormancy release and seedling emergence in Chinese cork oak. Using a thermal time model, we demonstrated for the first time that cold stratification reduces the heat requirement for seedling emergence. Physiological analyses revealed decreased starch content alongside increased soluble sugar levels, following cold stratification. Transcriptomic and proteomic data further indicated that genes associated with secondary metabolite biosynthesis and carbohydrate metabolism were significantly enriched, suggesting a metabolic shift that reduces heat requirements. Our findings provide a novel model for calculating the heat requirements of seedling emergence in Chinese cork oak and offer new insights into the role of cold stratification in optimizing seedling establishment. This study advances our understanding of the physiological and molecular mechanisms underlying epicotyl dormancy release and highlights the importance of cold stratification in epicotyl dormancy.
Due to the high value of spatio-temporal series forecasting, it has been regarded as a high-priority research topic in the fields of economics, physics, and transportation. Spatio-temporal graph neural networks can extract temporal and spatial information, thus suitable for spatio-temporal series prediction tasks. After the spatial information has been determined, the effectiveness of temporal correlation information extraction determines the accuracy of the prediction method. However, traditional methods cannot simultaneously balance computational consumption and long-range prediction capability in terms of temporal information. Therefore, we propose MAMGNN which contains a temporal information extraction (TIE) module that can ensure the accuracy of long series prediction at low computational consumption due to its structure. It also employs an inter-frame value variation trend constraint mechanism to ensure the accuracy of small-scale forecasts. Sufficient experiments demonstrate that our method achieves high prediction accuracy on economic and transportation datasets.
Objective:An applied problem facing all areas of data science is harmonizing data sources. Joining data from multiple origins with unmapped and only partially overlapping features is a prerequisite to developing and testing robust, generalizable algorithms, especially in healthcare. This integrating is usually resolved using meta-data such as feature names, which may be unavailable or ambiguous. Our goal is to design methods that create a mapping between structured tabular datasets derived from electronic health records independent of meta-data.Methods:We evaluate methods in the challenging case of numeric features without reliable and distinctive univariate summaries, such as nearly Gaussian and binary features. We assume that a small set of features are a priori mapped between two datasets, which share unknown identical features and possibly many unrelated features. Inter-feature relationships are the main source of identification which we expect. We compare the performance of contrastive learning methods for feature representations, novel partial auto-encoders, mutual-information graph optimizers, and simple statistical baselines on simulated data, public datasets, the MIMIC-III medical-record changeover, and perioperative records from before and after a medical-record system change. Performance was evaluated using both mapping of identical features and reconstruction accuracy of examples in the format of the other dataset.Results:Contrastive learning-based methods overall performed the best, often substantially beating the literature baseline in matching and reconstruction, especially in the more challenging real data experiments. Partial auto-encoder methods showed on-par matching with contrastive methods in all synthetic and some real datasets, along with good reconstruction. However, the statistical method we created performed reasonably well in many cases, with much less dependence on hyperparameter tuning. When validating feature match output in the EHR dataset we found that some mistakes were actually a surrogate or related feature as reviewed by two subject matter experts.Conclusion:In simulation studies and real-world examples, we find that inter-feature relationships are effective at identifying matching or closely related features across tabular datasets when meta-data is not available. Decoder architectures are also reasonably effective at imputing features without an exact match.
Hydrolyzable tannins (HTs) have garnered significant attention due to their proven beneficial effects in the clinical treatment of various diseases. The cupule of Chinese cork oak (Quercus variabilis Blume) has been used as raw material of traditional medicine for centuries for its high content of HTs. Previous studies have identified UGT84A13 as a key enzyme in the HT biosynthesis pathway in Q. variabilis, but the transcriptional regulation network of UGT84A13 remains obscure. Here, we performed a comprehensive genome-wide identification of the TCP transcription factors in Q. variabilis, elucidating their molecular evolution and gene structure. Gene expression analysis showed that TCP3 from the CIN subfamily and TCP6 from the PCF subfamily were co-expressed with UGT84A13 in cupule. Further functional characterization using dual-luciferase assays confirmed that TCP3, rather than TCP6, played a role in the transcriptional regulation of UGT84A13, thus promoting HT biosynthesis in the cupule of Q. variabilis. Our work identified TCP family members in Q. variabilis for the first time, and provided novel insights into the transcriptional regulatory network of UGT84A13 and HT biosynthesis in Q. variabilis, explaining the reason why the cupule enriches HTs that could be used for traditional medicine.
Inoculation of Lactiplantibacillus plantarum before yeast has the potential to affect the blueberry wine quality. The effects of four L. plantarum strains on the anthocyanins and volatile profile of blueberry wine were studied. Different L. plantarum inoculated wines presented a great color variation, and B3 strain was able to retain more bluish hue of blueberry wine. Compared with RF wine fermented by only Saccharomyces cerevisiae yeast, L. plantarum inoculated wines exhibited a better preservation of monomeric and acylated anthocyanins. More importantly, inoculation of L. plantarum, especially for B3 and B4 strains, enhanced the accumulation of pyranoanthocyanins, which may improve the color stability of blueberry wine. In addition, inoculation with L. plantarum favored the formation of lactate-related esters, acetoin, C6 alcohols, terpenes, aldehydes, and volatile phenols. Flash-Profile analysis indicated that B3, B4 and SS6 strains conferred wines more sour and astringent tastes, along with more ‘fruity’ and ‘chemical’ aromas.
BackgroundPost-operative complications present a challenge to the healthcare system due to the high unpredictability of their incidence. Socioeconomic conditions have been established as social determinants of health. However, their contribution relating to postoperative complications is still unclear as it can be heterogeneous based on community, type of surgical services, and sex and gender. Uncovering these relations can enable improved public health policy to reduce such complications.MethodsIn this study, we conducted a large population cross-sectional analysis of social vulnerability and the odds of various post-surgical complications. We collected electronic health records data from over 50,000 surgeries that happened between 2012 and 2018 at a quaternary health center in St. Louis, Missouri, United States and the corresponding zip code of the patients. We built statistical logistic regression models of postsurgical complications with the social vulnerability index of the tract consisting of the zip codes of the patient as the independent variable along with sex and race interaction.ResultsOur sample from the St. Louis area exhibited high variance in social vulnerability with notable rapid increase in vulnerability from the south west to the north of the Mississippi river indicating high levels of inequality. Our sample had more females than males, and females had slightly higher social vulnerability index. Postoperative complication incidence ranged from 0.75% to 41% with lower incidence rate among females. We found that social vulnerability was associated with abnormal heart rhythm with socioeconomic status and housing status being the main association factors. We also found associations of the interaction of social vulnerability and female sex with an increase in odds of heart attack and surgical wound infection. Those associations disappeared when controlling for general health and comorbidities.ConclusionsOur results indicate that social vulnerability measures such as socioeconomic status and housing conditions could affect postsurgical outcomes through preoperative health. This suggests that the domains of preventive medicine and public health should place social vulnerability as a priority to achieve better health outcomes of surgical interventions.
Emergence heterogeneity caused by epicotyl dormancy contributes to variations in seedling quality during largescale breeding. However, the mechanism of epicotyl dormancy release remains obscure. We first categorized the emergence stages of Chinese cork oak (Quercus variabilis) using the BBCH-scale. Subsequently, we identified the key stage of the epicotyl dormancy process. Our findings indicated that cold stratification significantly released epicotyl dormancy by increasing the levels of gibberellic acid 3 (GA3) and GA4. Genes associated with GA biosynthesis and signaling also exhibited altered expression patterns. Inhibition of GA biosynthesis by paclobutrazol (PAC) treatment severely inhibited emergence, with no effect on seed germination. Different concentrations (50 mu M, 100 mu M, and 200 mu M) of GA3 and GA4+7 treatments of germinated seeds demonstrated that both can promote the emergence, with GA4 exhibiting a more pronounced effect. In conclusion, this study provides valuable insights into the characterization of epicotyl dormancy in Chinese cork oak and highlights the critical role of GA biosynthesis in seedling emergence. These findings serve as a basis for further investigations on epicotyl dormancy and advancing large-scale breeding techniques.
Chinese cork oak (Quercus variabilis Blume) is a widespread tree species with high economic and ecological values. Chinese cork oak exhibits epicotyl dormancy, causing emergence heterogeneity and affecting the quality of seedling cultivation. Gibberellic acid-stimulated transcript (GAST) is a plant-specific protein family that plays a crucial regulatory role in plant growth, development, and seed germination. However, their evolution in Chinese cork oak and roles in epicotyl dormancy are still unclear. Here, a genome-wide identification of the GAST gene family was conducted in Chinese cork oak. Ten QvGAST genes were identified, and nine of them were expressed in seed. The physicochemical properties and promoter cis-acting elements of the selected Chinese cork oak GAST family genes indicated that the cis-acting elements in the GAST promoter are involved in plant development, hormone response, and stress response. Germinated seeds were subjected to gibberellins (GAs), abscisic acid (ABA), and fluridone treatments to show their response during epicotyl dormancy release. Significant changes in the expression of certain QvGAST genes were observed under different hormone treatments. QvGAST1, QvGAST2, QvGAST3, and QvGAST6 exhibited upregulation in response to gibberellin. QvGAST2 was markedly upregulated during the release of epicotyl dormancy in response to GA. These findings suggested that QvGAST2 might play an important role in epicotyl dormancy release. This study provides a basis for further analysis of the mechanisms underlying the alleviation of epicotyl dormancy in Chinese cork oak by QvGASTs genes.
Hydrolyzable tannins (HTs), predominant polyphenols in oaks, are widely used in grape wine aging, feed additives, and human healthcare. However, the limited availability of a high-quality reference genome of oaks greatly hampered the recognition of the mechanism of HT biosynthesis. Here, high-quality reference genomes of three Asian oak species (Quercus variabilis, Quercus aliena, and Quercus dentata) that have different HT contents were generated. Multi-omics studies were carried out to identify key genes regulating HT biosynthesis. In vitro enzyme activity assay was also conducted. Dual-luciferase and yeast one-hybrid assays were used to reveal the transcriptional regulation. Our results revealed that β-glucogallin was a biochemical marker for HT production in the cupules of the three Asian oaks. UGT84A13 was confirmed as the key enzyme for β-glucogallin biosynthesis. The differential expression of UGT84A13, rather than enzyme activity, was the main reason for different β-glucogallin and HT accumulation. Notably, sequence variations in UGT84A13 promoters led to different trans-activating activities of WRKY32/59, explaining the different expression patterns of UGT84A13 among the three species. Our findings provide three high-quality new reference genomes for oak trees and give new insights into different transcriptional regulation for understanding β-glucogallin and HT biosynthesis in closely related oak species.
Classifying birds accurately is essential for ecological monitoring. In recent years, bird image classification has become an emerging method for bird recognition. However, the bird image classification task needs to face the challenges of high intraclass variance and low inter-class variance among birds, as well as low model efficiency. In this paper, we propose a fine-grained bird classification method based on attention and decoupled knowledge distillation. First of all, we propose an attention-guided data augmentation method. Specifically, the method obtains images of the object's key part regions through attention. It enables the model to learn and distinguish fine features. At the same time, based on the localization-recognition method, the bird category is predicted using the object image with finer features, which reduces the influence of background noise. In addition, we propose a model compression method of decoupled knowledge distillation. We distill the target and nontarget class knowledge separately to eliminate the influence of the target class prediction results on the transfer of the nontarget class knowledge. This approach achieves efficient model compression. With 67% fewer parameters and only 1.2 G of computation, the model proposed in this paper still has a 87.6% success rate, while improving the model inference speed.
Tannins are useful for many industrial applications owing to their strong chelating, antibacterial, and antioxidant properties. Hydrolysable tannins (HTs) are suitable for feed additives to replace antibiotics because of their low anti-nutritional effects. However, regulation of HT biosynthesis in plants remains largely unknown. Here, we found that the HTs are predominant phenols in Chinese cork oak (Quercus variabilis) and their contents in cupules were much higher than those in other tissues; simultaneously, the expression of UGT84A13, a key gene for the first step of HT biosynthesis, was also high. RNA-seq analysis revealed 351 candidate genes whose expressions were correlated with UGT84A13 expressions and HT contents. Comprehensive analysis of the HD-Zip family identified subfamily I members—HB20 and HB26—as candidate genes responsible for regulation of UGT84A13 expression and HT biosynthesis. Subcellular localization assays showed that both HD-Zip proteins were localized to the nucleus. Interestingly, HB20 and HB26 had distinct trans-activating domains. Y1H and luciferase assays finally confirmed that HB26, instead of HB20, interacted with UGT84A13 promoter and activated its expression. This study provides valuable data for different tissues of Q. variabilis and new insights into the transcriptional regulation of UGT84A13 and HT biosynthesis, thereby laying the foundation for molecular breeding and enhancement of HT biosynthesis.
Real-time monitoring of microbial dynamics during fermentation is essential for wine quality control. This study developed a method that combines the fluorescent dye propidium monoazide (PMA) with CELL-qPCR, which can distinguish between dead and live microbes for Lactiplantibacillus plantarum . This method could detect the quantity of microbes efficiently and rapidly without DNA extraction during wine fermentation. The results showed that (1) the PMA-CELL-qPCR enumeration method developed for L. plantarum was optimized for PMA treatment concentration, PMA detection sensitivity and multiple conditions of sample pretreatment in wine environment, and the optimized method can accurately quantify 10 4 –10 8 CFU/mL of the target strain ( L. plantarum ) in multiple matrices; (2) when the concentration of dead bacteria in the system is 10 4 times higher than the concentration of live bacteria, there is an error of 0.5–1 lg CFU/mL in the detection results. The optimized sample pretreatment method in wine can effectively reduce the inhibitory components in the qPCR reaction system; (3) the optimized PMA-CELL-qPCR method was used to monitor the dynamic changes of L. plantarum during the fermentation of Cabernet Sauvignon wine, and the results were consistent with the plate counting method. In conclusion, the live bacteria quantification method developed in this study for PMA-CELL-qPCR in L. plantarum wines is accurate in quantification and simple in operation, and can be used as a means to accurately monitor microbial dynamics in wine and other fruit wines.
The use of neural networks for plant disease identification is a hot topic of current research. However, unlike the classification of ordinary objects, the features of plant diseases frequently vary, resulting in substantial intra-class variation; in addition, the complex environmental noise makes it more challenging for the model to categorize the diseases. In this paper, an attention and multidimensional feature fusion neural network (AMDFNet) is proposed for Camellia oleifera disease classification network based on multidimensional feature fusion and attentional mechanism, which improves the classification ability of the model by fusing features to each layer of the Inception structure and enhancing the fused features with attentional enhancement. The model was compared with the classical convolutional neural networks GoogLeNet, Inception V3, ResNet50, and DenseNet121 and the latest disease image classification network DICNN in a self-built camellia disease dataset. The experimental results show that the recognition accuracy of the new model reaches 86.78% under the same experimental conditions, which is 2.3% higher than that of GoogLeNet with a simple Inception structure, and the number of parameters is reduced to one-fourth compared to large models such as ResNet50. The method proposed in this paper can be run on mobile with higher identification accuracy and a smaller model parameter number.
We are currently in a network era which enables us to communicate more widely and more easily via the social networks. Meanwhile, negative information, such as fake news, rumors and computer viruses, often spread in social network. In order to restrain the propagation of such negative influence, we must find its sources in the network. But in real-world applications, we usually only know the scope of the negative influence spreading, and do not know who first propagates the negative influence. However, we can identify the sources of the negative influence based on the information of some observed nodes which are negatively influenced. This is the problem of influence sources locating. To tackle this problem, we present a latent space mapping-based method for identifying the multiple influence sources in the independent cascade model. The method first detects the candidate sources of the observed nodes based on message passing in a reversed network. An algorithm is presented to calculate the activation probability between nodes according to the influence spreading pattern in the independent cascade model. To evaluate each node’s rationality as the propagation source, we use the difference between the length of the path influencing an observed node and its activation time. We define two latent spaces, namely the influence senders and receivers’ latent spaces, and map the nodes into these two latent spaces to form a model describing the influence propagation. An estimation-maximization-based algorithm is proposed to optimize the propagation model. Based on this model, we propose a latent space mapping-based algorithm to identify the influence sources. The probability for each node to be a source is calculated by its positions in the latent spaces. Finally, k nodes with the largest probabilities are selected as the sources. Empirical results demonstrate that the influence sources identified by the proposed method can influence more observed nodes at more accurate time than other methods.
With the rapid growth of the internet, social networks provide an ideal platform for information exchange and propagation. Meanwhile, negative information, such as fake news, rumors, and computer viruses, often spread in social networks. To restrain the propagation of such negative information, we must find the sources of the negative influence. However, in real world applications, we usually only know the scope of the negative influence spreading and do not know who first propagates the negative influence. However, we can identify the sources of the negative influence based on the information of some observed nodes that are negatively influenced. We define this as the influencing source location problem. In this work, we present a network sparsification and stratification-based method to effectively locate multiple propagation sources using information from a few observed nodes. To reduce the complexity of the problem, we first sparsify the network by removing some edges that do not significantly impact the influence propagation to the observed nodes. We then define the stratified propagation graph where the nodes are divided into several levels according to their degrees and the paths leading to the observed nodes. We propose a method for constructing the stratified propagation graph and calculating the likelihoods of the nodes being the sources influencing the observed nodes. Then, k nodes with the maximum likelihoods are selected as the sources. Abundant experimental results show that the influence sources identified by the proposed method can influence more observed nodes at a more accurate time than other algorithms.
The problem with distance-aware influence maximization on multiple query locations (DIM-MQL) is selecting a group of nodes in the network to influence the nodes in the widest range possible near multiple query locations. A random walk-based algorithm for the DIM-MQL problem is presented. To accelerate query processing in real-time, our method involves offline and online processing. Offline processing conducts computations that are independent of the queries, and online processing answers queries in real-time. For offline processing, an algorithm is presented to estimate the upper and lower bounds of the influence spreading of the nodes based on a set of anchor points. We propose an algorithm to sample the influence spreading paths and estimate the influence spreading of the nodes. The number of samples required is analyzed and estimated. Based on the random walk approach, an algorithm is proposed to select anchor points by partitioning the nodes into groups. An algorithm is presented for seed selection in online processing. Based on the spreading bounds obtained in offline processing, a pruning technique is employed to accelerate query processing. Our empirical results show that the proposed algorithm can obtain a larger distance-aware influence spreading than other approaches.
The purpose of the experiment was to explore the localization and seasonal expression of extracellular signal regulated kinase (ERK) in the colonic tissue of wild ground squirrels (Spermophilus dauricus). Hematoxylin–eosin staining, immunohistochemistry, real-time quantitative PCR and Western blotting were used in this experiment. The histological results showed that the diameter of the colon lumen enlarged and the number of glandular cells increased in the non-breeding season. It was found in the immunochemical results that both ERK1/2 and pERK1/2 were expressed in the cytoplasm of goblet cells and intestinal epithelial cells, while pERK1/2 was also expressed in the nucleus of them. The immune localization of both was more obvious in the non-breeding season, especially in intestinal epithelial cells. Real-time quantitative PCR and Western blotting showed that ERK1/2 and pERK1/2 were seasonally highly expressed in the non-breeding season. The expression of ERK1/2 and pERK1/2 was seasonal changes and had significant increases in the non-breeding season. This study revealed that ERK1/2 had potential roles in the colon to the adaptation of seasonal changes in wild ground squirrels.