The National Health and Nutrition Examination Survey (NHANES) provides extensive public data on demographics, health, and nutrition, collected in 2-year cycles since 1999. Although invaluable for epidemiological and health-related research, the complexity of NHANES data, involving numerous files and disjoint metadata, makes accessing, managing, and analysing these datasets challenging. This paper presents a reproducible computational environment built upon Docker containers, PostgreSQL databases, and R/RStudio, designed to streamline NHANES data management, facilitate rigorous quality control, and simplify analyses across multiple survey cycles. We introduce specialized tools, such as the enhanced nhanesA R package and the phonto R package, to provide fast access to data, to help manage metadata, and to handle complexities arising from questionnaire design and cross-cycle data inconsistencies. Furthermore, we describe the Epiconnector platform, established to foster collaborative sharing of code, analytical scripts, and best practices, which taken together, can significantly enhance the reproducibility, extensibility, and robustness of scientific research using NHANES data.
The National Health and Nutrition Examination Survey provides comprehensive data on demographics, sociology, health and nutrition. Conducted in 2-year cycles since 1999, most of its data are publicly accessible, making it pivotal for research areas like studying social determinants of health or tracking trends in health metrics such as obesity or diabetes. Assembling the data and analyzing it presents a number of technical and analytic challenges. This paper introduces the nhanesA R package, which is designed to assist researchers in data retrieval and analysis and to enable the sharing and extension of prior research efforts. We believe that fostering community-driven activity in data reproducibility and sharing of analytic methods will greatly benefit the scientific community and propel scientific advancements.Database URL: https://github.com/cjendres1/nhanes
The human embryo derives from fusion of oocyte and sperm, undergoes growth and differentiation, resulting in a blastocyst. To initiate implantation, the blastocyst hatches from the zona pellucida, allowing access from external inputs. Modelling of uterine sperm distribution indicates that 200-5000 sperm cells may reach the implantation-stage blastocyst following natural coitus. We show ultrastructural evidence of sperm cells intruding into trophectoderm cells of zona-free blastocysts obtained from the uterus of rhesus monkeys. Interaction between additional sperm and zona-free blastocyst could be an evolutionary feature yielding adaptive processes influencing the developmental fate of embryos. This process bears potential implications in pregnancy success, sperm competition and human health.
Abstract The early human embryo derived from fusion of an oocyte with a single sperm undergoes growth and differentiation and results in an implantation-ready blastocyst. To initiate implantation, the blastocyst hatches from the zona pellucida, thus making it accessible to external inputs. Our modelling of sperm distribution through the uterus indicates that 200–5000 sperms following natural coitus during mid-luteal phase are in a position of reaching the implantation-stage blastocyst in the maternal uterus. We indeed have ultrastructural evidence of sperm cells intruding into the trophectoderm cells of uterine zona-free blastocysts obtained from rhesus monkeys. The question arises whether the negotiation between additional sperm and azonal blastocyst is a feature of evolution yielding adaptation processes influencing the developmental fate of an individual embryo or a neutral by-product in placental mammals. This process potentially bears implications in pregnancy success, sperm competition, and human health.
Plotting functions in R can be divided into three basic groups: High-level plotting functions create a new plot on the graphics device, possibly with axes, labels, titles and so on. Low-level plotting functions add more information to an existing plot, such as extra points, lines and labels. Interactive graphics functions allow you interactively add information to, or extract information from, an existing plot, using a pointing device such as a mouse. In addition, R maintains a list of graphical parameters that affect the result of various plot functions.
Using computer search algorithms, second order designs of composite type over a k-cube [-1, l](k) are obtained, where k is the number of factors. The advantages of the proposed approach are that (i) it is possible to obtain designs with higher D-efficiencies than a comparable orthogonal array composite design (OACD), and (ii) designs with fewer points than those required by an OACD and having comparable D-efficiencies can be obtained.
We construct a new family of orthogonal Latin hypercube designs having second order property with n rows and m = 4 columns, where n equivalent to 3 (mod 4). In particular, if n equivalent to 3 (mod 16), then we also report a family of such designs with m = 6 columns.
Random matrices whose entries come from a stationary Gaussian process are studied. The limiting behavior of the eigenvalues as the size of the matrix goes to infinity is the main subject of interest in this work. It is shown that the limiting spectral distribution is determined by the absolutely continuous component of the spectral measure of the stationary process. This is similar to the situation where the entries of the matrix are i.i.d. On the other hand, the discrete component contributes to the limiting behavior of the eigenvalues after a different scaling. Therefore, this helps to define a boundary between short and long range dependence of a stationary Gaussian process in the context of random matrices.
Abstract In trellis graphics one, two, or more variables are plotted and conditioned on different values or categories of a given variable (or variables). Component panels representing subsets of the data defined by levels of the conditioning variables are arranged in a single display in a manner that makes comparison easier. Many different types of plots lend themselves to this treatment, including scatterplots and three‐dimensional scatterplots. Trellis graphics help in understanding both the structure of the data and how well proposed models for the data actually fit.
Multiple myeloma (MM), a malignancy of plasma cells, is characterized by widespread genomic heterogeneity and, consequently, differences in disease progression and drug response. Although recent large-scale sequencing studies have greatly improved our understanding of MM genomes, our knowledge about genomic structural variation in MM is attenuated due to the limitations of commonly used sequencing approaches. In this study, we present the application of optical mapping, a single-molecule, whole-genome analysis system, to discover new structural variants in a primary MM genome. Through our analysis, we have identified and characterized widespread structural variation in this tumor genome. Additionally, we describe our efforts toward comprehensive characterization of genome structure and variation by integrating our findings from optical mapping with those from DNA sequencing-based genomic analysis. Finally, by studying this MM genome at two time points during tumor progression, we have demonstrated an increase in mutational burden with tumor progression at all length scales of variation.
In this article, we show the existence of limiting spectral distribution of a symmetric random matrix whose entries come from a stationary Gaussian process with covariances satisfying a summability condition. We provide an explicit description of the moments of the limiting measure. We also show that in some special cases the Gaussian assumption can be relaxed. The description of the limiting measure can also be made via its Stieltjes transform which is characterized as the solution of a functional equation. In two special cases, we get a description of the limiting measure - one as a free product convolution of two distributions, and the other one as a dilation of the Wigner semicircular law.
Latin hypercube designs have been found very useful for designing computer experiments. In recent years, several methods of constructing orthogonal Latin hypercube designs have been proposed in the literature. In this article, we report some more results on Latin hypercube designs, including several new designs.
Latin hypercube designs have been found very useful for designing computer experiments. In recent years, several methods of constructing orthogonal Latin hypercube designs have been proposed in the literature. In this article, we report some more results on the construction of orthogonal Latin hypercubes which result in several new designs.
Background Solid tumors present a panoply of genomic alterations, from single base changes to the gain or loss of entire chromosomes. Although aberrations at the two extremes of this spectrum are readily defined, comprehensive discernment of the complex and disperse mutational spectrum of cancer genomes remains a significant challenge for current genome analysis platforms. In this context, high throughput, single molecule platforms like Optical Mapping offer a unique perspective. Results Using measurements from large ensembles of individual DNA molecules, we have discovered genomic structural alterations in the solid tumor oligodendroglioma. Over a thousand structural variants were identified in each tumor sample, without any prior hypotheses, and often in genomic regions deemed intractable by other technologies. These findings were then validated by comprehensive comparisons to variants reported in external and internal databases , and by selected experimental corroborations. Alterations range in size from under 5 kb to hundreds of kilobases, and comprise insertions, deletions, inversions and compound events. Candidate mutations were scored at sub-genic resolution and unambiguously reveal structural details at aberrant loci. Conclusions The Optical Mapping system provides a rich description of the complex genomes of solid tumors, including sequence level aberrations, structural alterations and copy number variants that power generation of functional hypotheses for oligodendroglioma genetics.
The Optical Mapping System constructs ordered restriction maps spanning entire genomes through the assembly and analysis of large datasets comprising individually analyzed genomic DNA molecules. Such restriction maps uniquely reveal mammalian genome structure and variation, but also raise computational and statistical questions beyond those that have been solved in the analysis of smaller, microbial genomes. We address the problem of how to filter maps that align poorly to a reference genome. We obtain map-specific thresholds that control errors and improve iterative assembly. We also show how an optimal self-alignment score provides an accurate approximation to the probability of alignment, which is useful in applications seeking to identify structural genomic abnormalities.
Although microRNAs (miRNAs) are important regulators of gene expression, the transcriptional regulation of miRNAs themselves is not well understood. We employed an integrative computational pipeline to dissect the transcription factors (TFs) responsible for altered miRNA expression in ovarian carcinoma. Using experimental data and computational predictions to define miRNA promoters across the human genome, we identified TFs with binding sites significantly overrepresented among miRNA genes over-expressed in ovarian carcinoma. This pipeline nominated TFs of the p53/p63/p73 family as candidate drivers of miRNA overexpression. Analysis of data from an independent set of 253 ovarian carcinomas in The Cancer Genome Atlas showed that p73 and p63 expression is significantly correlated with expression of miRNAs whose promoters contain p53/p63/p73 family binding sites. In experimental validation of specific miRNAs predicted by the analysis to be regulated by p73 and p63, we found that p53/p63/p73 family binding sites modulate promoter activity of miRNAs of the miR-200 family, which are known regulators of cancer stem cells and epithelial-mesenchymal transitions. Furthermore, in chromatin immunoprecipitation studies both p73 and p63 directly associated with the miR-200b/a/429 promoter. This study delineates an integrative approach that can be applied to discover transcriptional regulatory mechanisms in other biological settings where analogous genomic data are available.