PURPOSE:To investigate the impact of vitamin D deficiency (VDD) on retinal and choroidal thickness (CT) in patients with diabetes mellitus (DM) without diabetic retinopathy (DR) using swept-source optical coherence tomography (SS-OCT). METHODS:Fifty-six eyes from 56 DM-NDR patients were included and divided into two groups based on serum 25(OH)D levels: VDD Group (25(OH)D ≤ 20 ng/mL, n = 27) and non-VDD Group (25(OH)D > 20 ng/mL, n = 29). Macular OCT imaging (6 × 6 mm) was performed and divided into the foveal region (0-1 mm), parafoveal region (1-3 mm), and perifoveal region (3-6 mm). The latter two zones were further subdivided into four quadrants: temporal, superior, nasal, and inferior (T, S, N, I). Retinal layer (RL), retinal nerve fiber layer (RNFL), ganglion cell layer plus inner plexiform layer (GCL-IPL), and CT were measured. Serum vitamin D levels were quantified using electrochemiluminescence immunoassay (ECLIA). RESULTS:Significant differences in RL and GCL-IPL thicknesses were observed within the foveal region, and in RNFL thickness across all quadrants in the parafoveal region (p < 0.05). RL thickness showed significant differences in the superior and temporal quadrants (p < 0.05). GCL-IPL thickness showed significant differences in all quadrants except the inferior quadrant (p < 0.05). CONCLUSION:VDD is associated with reduced RL, RNFL and GCL-IPL thicknesses, particularly in the parafoveal region, indicating that these layers may serve as sensitive biomarkers for VDD-induced neurodegeneration in diabetic patients. Vitamin D deficiency may be a risk factor for the progression of DR.
RNA N4-acetylcytidine (ac4C) is the acetylation of cytidine at the nitrogen-4 position, which is a highly conserved RNA modification and involves a variety of biological processes. Hence, accurate identification of genome-wide ac4C sites is vital for understanding regulation mechanism of gene expression. In this work, a novel predictor, named iRNA-ac4C, was established to identify ac4C sites in human mRNA based on three feature extraction methods, including nucleotide composition, nucleotide chemical property, and accumulated nucleotide frequency. Subsequently, minimum-Redundancy-Maximum-Relevance combined with incremental feature selection strategies was utilized to select the optimal feature subset. According to the optimal feature subset, the best ac4C classification model was trained by gradient boosting decision tree with 10-fold cross-validation. The results of independent testing set indicated that our proposed method could produce encouraging generalization capabilities. For the convenience of other researchers, we established a user-friendly web server which is freely available at http://lin-group.cn/server/iRNA-ac4C/. We hope that the tool could provide guide for wet-experimental scholars.
Large-scale screening for the risk of coronary heart disease (CHD) is crucial for its prevention and management. Physical examination data has the advantages of wide coverage, large capacity, and easy collection. Therefore, here we report a gender-specific cascading system for risk assessment of CHD based on physical examination data. The dataset consists of 39,538 CHD patients and 640,465 healthy individuals from the Luzhou Health Commission in Sichuan, China. Fifty physical examination characteristics were considered, and after feature screening, ten risk factors were identified. To facilitate large-scale CHD risk screening, a CHD risk model was developed using a fully connected network (FCN). For males, the model achieves AUCs of 0.8671 and 0.8659, respectively on the independent test set and the external validation set. For females, the AUCs of the model are 0.8991 and 0.9006, respectively on the independent test set and the external validation set. Furthermore, to enhance the convenience and flexibility of the model in clinical and real-life scenarios, we established a CHD risk scorecard base on logistic regression (LR). The results show that, for both males and females, the AUCs of the scorecard on the independent test set and the external verification set are only slightly lower (<0.05) than those of the corresponding prediction model, indicating that the scorecard construction does not result in a significant loss of information. To promote CHD personal lifestyle management, an online CHD risk assessment system has been established, which can be freely accessed at http://lin-group.cn/server/CHD/index.html .
Oxidative stress exerts a significant influence on the pathogenesis of various cataracts by inducing degradation and aggregation of lens proteins and apoptosis of lens epithelial cells. Keratinocyte growth factor−2 (KGF-2) exerts a favorable cytoprotective effect against oxidative stress in vivo and in vitro. In this work, we investigated the molecular mechanisms of KGF-2 against hydrogen peroxide- (H2O2-) induced oxidative stress and apoptosis in human lens epithelial cells (HLECs) and rat lenses. KGF-2 pretreatment could reduce H2O2-induced cytotoxicity as well as reactive oxygen species (ROS) accumulation. KGF-2 also increases B-cell lymphoma-2 (Bcl-2), quinine oxidoreductase-1 (NQO-1), superoxide dismutase (SOD2), and catalase (CAT) levels while decreasing the expression level of Bcl2-associated X (Bax) and cleaved caspase-3 in H2O2-stimulated HLECs. LY294002, the phosphatidylinositol-3-kinase (PI3K)/Akt inhibitor, abolished KGF-2's effect to some extent, demonstrating that KGF-2 protected HLECs via the PI3K/Akt pathway. On the other hand, KGF-2 activated the Nrf2/HO-1 pathway by regulating the PI3K/Akt pathway. Silencing nuclear factor erythroid 2-related factor 2 (Nrf2) by targeted-siRNA and inhibiting heme oxygenase-1 (HO-1) through zinc protoporphyrin IX (ZnPP) significantly decreased cytoprotection of KGF-2. Furthermore, as revealed by lens organ culture assays, KGF-2 treatment decreased H2O2-induced lens opacity in a concentration-dependent manner. As demonstrated by these data, KGF-2 resisted H2O2-mediated apoptosis and oxidative stress in HLECs through Nrf2/HO-1 and PI3K/Akt pathways, suggesting a potential protective effect against the formation of cataracts.
With the rapid development of science and technology, the trend of low age myopia is becoming increasingly significant. The latest national survey done by the Chinese government found that more than 80% of Chinese teenagers suffer from myopia. Adolescent myopia is closely related to living environment, heredity, and living habits. Quantifying the relationship between myopia and living environment, heredity, and living habits is conductive to the prevention and intervention of adolescent myopia. In this study, we investigated the relationships between four main factors (environment, habits, parental vision, and demographic) and myopia status by analyzing the questionnaire data. Data were collected from Chengdu, China in 2021, including 2808 myopia samples and 5693 non-myopia samples, with a total of 22 features. Then, these 22 features were inputted into three machine learning algorithms to discriminate the two classes of samples. Results show that the computational model could produce an AUC of 0.768. To pick out the most important features which play important roles in classification, we used incremental feature selection strategy to screen the 22 features. As a result, we found that the 4 most influential features with XGBoost could achieve a competitive AUC of 0.764. To further investigate the risk and protective factors affecting adolescent myopia, we used OR values derived from MLE-LR to analyze the relationship between 22 features and adolescent myopia. Results showed that the age variable was the most significant risk factor for myopia, followed by the myopia status of parents. The most protective factor for eyesight is the measure taken by the children, followed by the distance between books and eyes when reading. These discoveries can guide the prevention and control of myopia in children and adolescents.
Diabetes is a metabolic disorder caused by insufficient insulin secretion and insulin secretion disorders. From health to diabetes, there are generally three stages: health, pre-diabetes and type 2 diabetes. Early diagnosis of diabetes is the most effective way to prevent and control diabetes and its complications. In this work, we collected the physical examination data from Beijing Physical Examination Center from January 2006 to December 2017, and divided the population into three groups according to the WHO (1999) Diabetes Diagnostic Standards: normal fasting plasma glucose (NFG) (FPG < 6.1 mmol/L), mildly impaired fasting plasma glucose (IFG) (6.1 mmol/L ≤ FPG < 7.0 mmol/L) and type 2 diabetes (T2DM) (FPG > 7.0 mmol/L). Finally, we obtained1,221,598 NFG samples, 285,965 IFG samples and 387,076 T2DM samples, with a total of 15 physical examination indexes. Furthermore, taking eXtreme Gradient Boosting (XGBoost), random forest (RF), Logistic Regression (LR), and Fully connected neural network (FCN) as classifiers, four models were constructed to distinguish NFG, IFG and T2DM. The comparison results show that XGBoost has the best performance, with AUC (macro) of 0.7874 and AUC (micro) of 0.8633. In addition, based on the XGBoost classifier, three binary classification models were also established to discriminate NFG from IFG, NFG from T2DM, IFG from T2DM. On the independent dataset, the AUCs were 0.7808, 0.8687, 0.7067, respectively. Finally, we analyzed the importance of the features and identified the risk factors associated with diabetes.
TAPS uses a standard modificationcalling tool called asTair (H.
OBJECTIVE:Early diagnosis of diabetic kidney disease (DKD) has long been a complex problem. This study aimed to analyze the metabolomic characteristics of plasma extracellular vesicles (EVs) at different stages of DKD in order to evaluate the metabolites of plasma EVs and select new biomarkers for the early diagnosis of DKD.PATIENTS AND METHODS:A total of 78 plasma samples were collected, including samples from 20 healthy controls, 20 patients with type 2 diabetes mellitus (T2DM), 18 patients with DKD stage III, and 20 patients with DKD stage IV. In addition, EVs were isolated for metabolomics analysis.RESULTS:The results identified differences in EV metabolomic characteristics in DKD patients at different stages, as well as significant differences in EV metabolomics between T2DM patients without DKD and patients with DKD. Ten Significantly differential metabolites were associated with the occurrence and progression of DKD. Uracil, LPC(O-18:1/0:0), sphingosine 1-phosphate, and 4-acetamidobutyric acid were identified as potential early biomarkers for DKD, showing excellent predictive performance.CONCLUSION:Uracil, LPC(O-18:1/0:0), sphingosine 1-phosphate, and 4-acetamidobutyric acid exhibited potential as suitable biomarkers for early DKD diagnosis. Unexpectedly, combining these four candidate metabolites resulted in enhanced predictive ability for DKD.
DNA modification plays a pivotal role in regulating gene expression in cell development. As prevalent markers on DNA, 5-methylcytosine (5mC), N6-methyladenine (6mA), and N4-methylcytosine (4mC) can be recognized by specific methyltransferases, facilitating cellular defense and the versatile regulation of gene expression in eukaryotes and prokaryotes. Recent advances in DNA sequencing technology have permitted the positions of different modifications to be resolved at the genome-wide scale, which has led to the discovery of several novel insights into the complexity and functions of multiple methylations. In this review, we summarize differences in the various mapping approaches and discuss their pros and cons with respect to their relative read depths, speeds, and costs. We also discuss the development of future sequencing technologies and strategies for improving the detection resolution of current sequencing technologies. Lastly, we speculate on the potentially instrumental role that these sequencing technologies might play in future research.
As a key region, promoter plays a key role in transcription regulation. A eukaryotic promoter database called EPD has been constructed to store eukaryotic POL II promoters. Although there are some promoter databases for specific prokaryotic species or specific promoter type, such as RegulonDB for Escherichia coli K-12, DBTBS for Bacillus subtilis and Pro54DB for sigma 54 promoter, because of the diversity of prokaryotes and the development of sequencing technology, huge amounts of prokaryotic promoters are scattered in numerous published articles, which is inconvenient for researchers to explore the process of gene regulation in prokaryotes. In this study, we constructed a Prokaryotic Promoter Database (PPD), which records the experimentally validated promoters in prokaryotes, from published articles. Up to now, PPD has stored 129,148 promoters across 63 prokaryotic species manually extracted from published papers. We provided a friendly interface for users to browse, search, blast, visualize, submit and download data. The PPD will provide relatively comprehensive resources of prokaryotic promoter for the study of prokaryotic gene transcription. The PPD is freely available and easy accessed at http://lin-group.cn/database/ppd/.
Motivation Protein carbonylation is one of the most important oxidative stress-induced post-translational modifications, which is generally characterized as stability, irreversibility and relative early formation. It plays a significant role in orchestrating various biological processes and has been already demonstrated to be related to many diseases. However, the experimental technologies for carbonylation sites identification are not only costly and time consuming, but also unable of processing a large number of proteins at a time. Thus, rapidly and effectively identifying carbonylation sites by computational methods will provide key clues for the analysis of occurrence and development of diseases. Results In this study, we developed a predictor called iCarPS to identify carbonylation sites based on sequence information. A novel feature encoding scheme called residues conical coordinates combined with their physicochemical properties was proposed to formulate carbonylated protein and non-carbonylated protein samples. To remove potential redundant features and improve the prediction performance, a feature selection technique was used. The accuracy and robustness of iCarPS were proved by experiments on training and independent datasets. Comparison with other published methods demonstrated that the proposed method is powerful and could provide powerful performance for carbonylation sites identification. Availability and implementation Based on the proposed model, a user-friendly webserver and a software package were constructed, which can be freely accessed at http://lin-group.cn/server/iCarPS. Supplementary information Supplementary data are available at Bioinformatics online.
Diabetes is a global epidemic. Long-term exposure to hyperglycemia can cause chronic damage to various tissues. Thus, early diagnosis of diabetes is crucial. In this study, we designed a computational system to predict diabetes risk by fusing multifarious types of physical examination data. We collected 1,507,563 physical examination data of healthy people and diabetes patients, as well as 387,076 physical examination data from the follow-up records from 2011 to 2017 of diabetes patients in Luzhou City in China. Three types of physical examination indexes were statistically analyzed: demographics, vital signs, and laboratory values. To distinguish diabetes patients from healthy people, a model based on eXtreme Gradient Boosting (XGBoost) was developed, which could produce an area under the receiver operating characteristic curve (AUC) of 0.8768. Moreover, to improve the convenience and flexibility of the model in clinical and real-life scenarios, a diabetes risk scorecard was established based on logistic regression, which could evaluate human health. Lastly, we statistically analyzed the data from the follow-up records to identify the key factors influencing patient control of their conditions. To improve the diabetes cascade screening and personal lifestyle management, an online diabetes risk assessment system was established, which can be freely accessed at http://lin-group.cn/server/DRSC/index.html. This system is expected to provide guidance for human health management.
DNase I hypersensitive sites (DHSs) are special regions of the chromosome with loose structure that can be recognized, bound, and cleaved by DNase I enzyme. In these specific regions, the chromatin lacks condensed structure, resulting in increased accessibility. DHSs are hallmarks of gene expression regulation, and the characterization of DHSs is important to understand transcriptional regulatory mechanism and also to facilitate localization of cis-regulatory elements such as promoters, enhancers, insulators, silencers, and locus control regions. Although many experimental methods have been proposed to identify DHSs, these methods are time-consuming and expensive, making it urgent to develop computational methods to predict DHSs. In this study, we described a sequence-based predictor to identity DHSs in the human genome. In the predictor, optimal features were selected from a large feature set including various k-mer nucleotide compositions and correlation information of physicochemical properties of dinucleotides by using a two-step feature selection algorithm. Using 5-fold cross-validation, the proposed method achieved a Matthews correlation coefficient and accuracy of 0.66 and 0.87, respectively, which are higher than those of published DHS predictors, indicating the good performance of our method. The benchmark datasets and trained DHS model are available at https://github.com/Jackie-Suv/ iDHS-SVM.
Allergens have the ability to enter the body and cause illness. Leukotriene is the widespread allergen which could stimulate mast cells to discharge histamine which causes allergy symptoms. An effective strategy for treating leukotriene-induced allergy is to find the inhibitors of leukotriene or histamine activity from phytochemicals. For this purpose, a library of 8,500 phytochemicals was generated using MOE software. The structures of histamine-1 receptor and cysteinyl leukotriene receptor-1 were predicted by the homology modeling method through the SWISS model. The phytochemicals were docked with predicted structures of histamine-1 and cysteinyl leukotriene receptor-1 in MOE software to determine the binding affinity of the phytochemicals against the targets. Moreover, chemoinformatics properties and ADMET of phytochemicals were assessed to find the drug likeness behavior of compounds. Compound ID 10054216 has the lowest S-score value for H-1 receptor that is -18.9186 kcal/mol which is lower than the value of standard -15.167 kcal/mol. The other compounds 393471, 71448939, 10722577, and 442614 also showed good S-score values than the standard. Moreover, compound ID 11843082 has the lowest S-score value for CL1R that is -15.481 kcal/mol which is lower than the value of standard -12.453 kcal/mol. The other compounds 72284, 5282102, 66559251, and 102506430 also showed good S-score values than the standard. In this research article, we performed molecular docking to find the best inhibitors against H1R and CL1R and their antiallergic efficacy. This in silico knowledge will be helpful in near future for the design of novel, safe, and less costing H-1 receptor and CL1R inhibitors with the aim to improve human life quality.
Background: In view of a lack of consensus regarding the biomarkers of pre-diabetes and type 2 diabetes mellitus (T2DM), we sought to develop integrated biomarker profiling (IBP) of the metabolome.Methods: We studied 1,705 subjects including normal glucose tolerance, impaired fasting glucose (IFG), T2DM, and hyperlipidemia at five clinical centers in China. We aimed to construct the IBPs of IFG and T2DM, designed discovery, test, and validation phases. The discovery and test phases were employed to identify the potential serum biomarkers for IFG and T2DM. In the validation phase, the potential biomarkers were screened using Gini impurity to construct the IBP of IFG and T2DM based on the eXtreme Gradient Boosting (XGBoost) model, and individuals with hyperlipidemia were set as an interference group to evaluate the diagnostic accuracy of the IBP.Findings: Forty-one potential biomarkers were identified, in which 16 potential biomarkers were quantitatively analyzed. The IBP, which holistically reflected a mutually correlated bio-network, was constructed based on the XGBoost (AUC=0·823) model and Gini impurity (Top 10), and consisted of lysophosphatidylcholine (P-16:0), L-isoleucine, L-arginine, L-carnitine, L-phenylalanine, L-glutamic acid, L-lysine, L-methionine, L-leucine, and acetyl-L-carnitine. Besides, the service website ( http://pdm.lin-group.cn/) of the IBPs for IFG and T2DM was established to serve the public.Interpretation: IBP of the metabolome provides more consolidated evidence than isolated biomarkers, and might improve auxiliary diagnostic value in IFG and T2DM.Trial Registration: Chinese Clinical Trial Registry (ChiCTR1800014301).Funding Statement: This work was supported by the National Natural Science Foundation of China (grant number 81773891, 82073741), the National Great New Drugs Development Project of China (grant number 2017ZX09301-040), the National Key R&D Program of China (grant number 2020YFF01014606), and the Beijing Excellent Talent Project (grant number DFL20190702, 2018000021223TD09). Declaration of Interests: We declare no competing interests. Ethics Approval Statement: The study protocols were approved by the Scientific Research Ethics Committee of Capital Medical University Affiliated Beijing Shijitan Hospital (2017-035) and registered on the Chinese Clinical Trial Registry (ChiCTR1800014301). All the participants provided their written informed consent.
Background: SARS-Cov-2 is a newly emerged coronavirus and causes a severe type of pneumonia in the host organism. So, it is an urgent need to find some inhibitors against SARS-Cov-2. Therefore, drug repurposing study is an effective strategy for treating pneumonia to find the inhibitors of SARS-Cov-2 proteins. Method: For this purpose, a library of 2500 verified drug chemical compounds was generated and the compounds were docked against Nucleocapsid, Membrane and Envelope protein structures of SARSCov- 2 to determine the binding affinity of the chemical compounds against targeting binding pockets. Moreover, cheminformatics properties and ADMET of these compounds were assessed to find the druglikeness behavior of compounds. The chemical compounds with the lowest S-score were identified as potential inhibitors. Results: Our findings showed that the compound ids 1212, 1019 and 1992 could interact inside the active sites of membrane protein, nucleocapsid protein and envelope protein. Conclusion: This in silico knowledge will be helpful for the design of novel, safe and less expensive drugs against the SARS-Cov-2.
Transcription factors play key roles in cell-fate decisions by regulating 3D genome conformation and gene expression. The traditional view is that methylation of DNA hinders transcription factors binding to them, but recent research has shown that many transcription factors prefer to bind to methylated DNA. Therefore, identifying such transcription factors and understanding their functions is a stepping-stone for studying methylation-mediated biological processes. In this paper, a two-step discriminated method was proposed to recognize transcription factors and their preference for methylated DNA based only on sequences information. In the first step, the proposed model was used to discriminate transcription factors from non-transcription factors. The areas under the curve (AUCs) are 0.9183 and 0.9116, respectively, for the 5-fold cross-validation test and independent dataset test. Subsequently, for the classification of transcription factors that prefer methylated DNA and transcription factors that prefer non-methylated DNA, our model could produce the AUCs of 0.7744 and 0.7356, respectively, for the 5-fold cross-validation test and independent dataset test. Based on the proposed model, a user-friendly web server called TFPred was built, which can be freely accessed at http://lin-group.cn/server/TFPred/.
Human visual acuity is anatomically determined by the retinal fovea. The ontogenetic development of the fovea can be seriously hindered by oculocutaneous albinism (OCA), which is characterized by a disorder of melanin synthesis. Although people of all ethnic backgrounds can be affected, no efficient treatments for OCA have been developed thus far, due partly to the lack of effective animal models. Rhesus macaques are genetically homologous to humans and, most importantly, exhibit structures of the macula and fovea that are similar to those of humans; thus, rhesus macaques present special advantages in the modeling and study of human macular and foveal diseases. In this study, we identified rhesus macaque models with clinical characteristics consistent with those of OCA patients according to observations of ocular behavior, fundus examination, and optical coherence tomography. Genomic sequencing revealed a biallelic p.L312I mutation in TYR and a homozygous p.S788L mutation in OCA2, both of which were further confirmed to affect melanin biosynthesis via in vitro assays. These rhesus macaque models of OCA will be useful animal resources for studying foveal development and for preclinical trials of new therapies for OCA.
The locations of the initiation of genomic DNA replication are defined as origins of replication sites (ORIs), which regulate the onset of DNA replication and play significant roles in the DNA replication process. The study of ORIs is essential for understanding the cell-division cycle and gene expression regulation. Accurate identification of ORIs will provide important clues for DNA replication research and drug development by developing computational methods. In this paper, the first integrated predictor named iORI-Euk was built to identify ORIs in multiple eukaryotes and multiple cell types. In the predictor, seven eukaryotic (Homo sapiens, Mus musculus, Drosophila melanogaster, Arabidopsis thaliana, Pichia pastoris, Schizosaccharomyces pombe and Kluyveromyces lactis) ORI data was collected from public database to construct benchmark datasets. Subsequently, three feature extraction strategies which are k-mer, binary encoding and combination of k-mer and binary were used to formulate DNA sequence samples. We also compared the different classification algorithms' performance. As a result, the best results were obtained by using support vector machine in 5-fold cross-validation test and independent dataset test. Based on the optimal model, an online web server called iORI-Euk (http://lin-group.cn/server/iORI-Euk/) was established for the novel ORI identification.
5hmC, 6mA, and 4mC are three common DNA modifications and are involved in various of biological processes. Accurate genome-wide identification of these sites is invaluable for better understanding their biological functions. Owing to the labor-intensive and expensive nature of experimental methods, it is urgent to develop computational methods for the genome-wide detection of these sites. Keeping this in mind, the current study was devoted to construct a computational method to identify 5hmC, 6mA, and 4mC. We initially used K-tuple nucleotide component, nucleotide chemical property and nucleotide frequency, and mono-nucleotide binary encoding scheme to formulate samples. Subsequently, random forest was utilized to identify 5hmC, 6mA, and 4mC sites. Cross-validated results showed that the proposed method could produce the excellent generalization ability in the identification of the three modification sites. Based on the proposed model, a web-server called iDNA-MS was established and is freely accessible at http://lin-group.cn/server/iDNA-MS.