Large language models (LLMs) achieve remarkable generative performance, yet their output quality is dependent on the decoding strategy. While sampling-based methods (e.g., top-k, nucleus) and search-and-select based methods (e.g., beam search, best-of-n, majority voting) can improve upon greedy decoding, both approaches suffer from limitations: sampling generally commits to a single path, while search often expends excessive computation regardless of task complexity. To address these, we introduce Entropy-informed decoding (EDEN), a plug-and-play, model-agnostic decoding framework that adaptively allocates computation based on the model's own uncertainty, approximating higher-width beam search with fewer expansions. At each generation step, EDEN estimates the entropy of the output token distribution and adjusts the branching factor monotonically with the entropy, expanding more candidates in high-entropy regions and following a greedier path in low-entropy regions, improving token efficiency. Experiments across complex tasks, including mathematical reasoning, code generation, and scientific questions, demonstrate that EDEN consistently improves output quality over existing decoding strategies, achieving better accuracy-expansion trade-offs than fixed-width beam search. By treating next-token selection as a noisy maximisation problem, we prove that branching factors monotone in entropy are guaranteed to find better (i.e. more probable) continuations than any fixed branching factor within the same total expansion budget, and derive explicit regret rates characterising the benefit of the adaptive allocation.
ABSTRACT Changes in the spatial organization of bacterial chromosomes under stress conditions and its biological implications remain poorly understood. We mapped the structural landscape of wild-type and Δ dcm E. coli chromosomes under triclosan stress using Hi-C to identify triclosan-induced chromosomal interaction domains (CIDs). Two CIDs were common to the wild-type and Δ dcm E. coli , including a CID with a common boundary at fabI gene, which encodes the triclosan target. All mutations and structural variants under triclosan stress were observed within or in close proximity to triclosan-induced CIDs. Absence of Dcm methylation impacts both short- and long-range interactions in triclosan stress. Single-base resolution methylome maps reveal hypermethylation of adenines (in wild-type and Δ dcm ) and cytosines (in wild-type) in the two common triclosan-induced CIDs. Furthermore, global gene expression profiling identified enrichment of highly expressed genes within the two common CIDs. Our findings suggest that stress-induced CIDs in E. coli are hotspots for genetic variations and are associated with enhanced transcriptional activity and hypermethylation of Dam/Dcm motifs.
Experimental evolution studies have examined coevolutionary dynamics between bacteria and lytic phages, where two models for antagonistic coevolution dominate: arms-race dynamics (ARD) and fluctuating-selection dynamics (FSD). Here, we tested the ability for Pseudomonas aeruginosa to coevolve with phage OMKO1 during 10 passages in the laboratory, whether ARD versus FSD coevolution occurred, and how coevolution affected a predicted phenotypic trade-off between phage resistance and antibiotic sensitivity. We used a unique "deep" sampling design, where 96 bacterial clones per passage were obtained from the three replicate coevolving communities. Next, we examined phenotypic changes in growth ability, susceptibility to phage infection and resistance to antibiotics. Results confirmed that the bacteria and phages coexisted throughout the study with one community undergoing ARD, whereas the other two showed evidence for FSD. Surprisingly, only the ARD bacteria demonstrated the anticipated trade-off. Whole genome sequencing revealed that treatment populations of bacteria accrued more de novo mutations, relative to a control bacterial population. Additionally, coevolved bacteria presented mutations in genes for biosynthesis of flagella, type-IV pilus and lipopolysaccharide, with three mutations fixing contemporaneously with the occurrence of the phenotypic trade-off in the ARD-coevolved bacteria. Our study demonstrates that both ARD and FSD coevolution outcomes are possible in a single interacting bacteria-phage system and that occurrence of predicted phage-driven evolutionary trade-offs may depend on the genetics underlying evolution of phage resistance in bacteria. These results are relevant for the ongoing development of lytic phages, such as OMKO1, in personalized treatment of human patients, as an alternative to antibiotics.
Open OnDemand (OOD) greatly lowers the barrier to entry to high performance computing (HPC) resources and facilitates usage for new and experienced users. Moreover, using OOD for courses with computational components enables lecturers and teaching assistants to focus on the course instead of on-boarding students. To these ends the Yale Center for Research Computing (YCRC) adopted OOD to make its advanced cyberinfrastructure more accessible to researchers and students across the university. While the use of single sign-on authentication in OOD makes HPC clusters easier to access, as implemented this authentication method limits users to a single HPC account per university-provided identity. The YCRC provides separate and temporary course-specific accounts to isolate course usage, which presented a challenge for supporting users with pre-existing research allocations or who are participating in multiple courses. In this paper, we present an easy to implement and maintain solution that separates traffic for each course and maps a single user identity to multiple HPC accounts. This solution can be used for other situations other than courses; wherever one single sign-on identity should access multiple HPC accounts.
OBJECTIVES:To measure the variability in carbapenem susceptibility conferred by different OxaAb variants, characterize the molecular evolution of oxaAb and elucidate the contribution of OxaAb and other possible carbapenem resistance factors in the clinical isolates using WGS and LC-MS/MS. METHODS:Antimicrobial susceptibility tests were performed on 10 clinical Acinetobacter baumannii isolates. Carbapenem MICs were evaluated for all oxaAb variants cloned into A. baumannii CIP70.10 and BM4547, with and without their natural promoters. Molecular evolution analysis of the oxaAb variants was performed using FastTree and SplitsTree4. Resistance determinants were studied in the clinical isolates using WGS and LC-MS/MS. RESULTS:Only the OxaAb variants with I129L and L167V substitutions, OxaAb(82), OxaAb(83), OxaAb(107) and OxaAb(110) increased carbapenem MICs when expressed in susceptible A. baumannii backgrounds without an upstream IS element. Carbapenem resistance was conferred with the addition of their natural upstream ISAba1 promoter. LC-MS/MS analysis on the original clinical isolates confirmed overexpression of the four I129L and L167V variants. No other differences in expression levels of proteins commonly associated with carbapenem resistance were detected. CONCLUSIONS:Elevated carbapenem MICs were observed by expression of OxaAb variants carrying clinically prevalent substitutions I129L and L167V. To drive carbapenem resistance, these variants required overexpression by their upstream ISAba1 promoter. This study clearly demonstrates that a combination of IS-driven overexpression of oxaAb and the presence of particular amino acid substitutions in the active site to improve carbapenem capture is key in conferring carbapenem resistance in A. baumannii and other mechanisms are not required.
Identifying our most distant animal relatives has emerged as one of the most challenging problems in phylogenetics. This debate has major implications for our understanding of the origin of multicellular animals and of the earliest events in animal evolution, including the origin of the nervous system. Some analyses identify sponges as our most distant animal relatives (Porifera-sister hypothesis), and others identify comb jellies (Ctenophora-sister hypothesis). These analyses vary in many respects, making it difficult to interpret previous tests of these hypotheses. To gain insight into why different studies yield different results, an important next step in the ongoing debate, we systematically test these hypotheses by synthesizing 15 previous phylogenomic studies and performing new standardized analyses under consistent conditions with additional models. We find that Ctenophora-sister is recovered across the full range of examined conditions, and Porifera-sister is recovered in some analyses under narrow conditions when most outgroups are excluded and site-heterogeneous CAT models are used. We additionally find that the number of categories in site-heterogeneous models is sufficient to explain the Porifera-sister results. Furthermore, our cross-validation analyses show CAT models that recover Porifera-sister have hundreds of additional categories and fail to fit significantly better than site-heterogenuous models with far fewer categories. Systematic and standardized testing of diverse phylogenetic models suggests that we should be skeptical of Porifera-sister results both because they are recovered under such narrow conditions and because the models in these conditions fit the data no better than other models that recover Ctenophora-sister.
A common claim of evolutionary computation methods is that they can achieve good results without the need for human intervention. However, one criticism of this is that there are still hyperparameters which must be tuned in order to achieve good performance. In this work, we propose a near parameter-free genetic programming approach, which adapts the hyperparameter values throughout evolution without ever needing to be specified manually. We apply this to the area of automated machine learning (by extending TPOT), to produce pipelines which can effectively be claimed to be free from human input, and show that the results are competitive with existing state-of-the-art which use hand-selected hyperparameter values. Pipelines begin with a randomly chosen estimator and evolve to competitive pipelines automatically. This work moves towards a truly automated approach to AutoML.
In an attempt to control the mosquito-borne diseases yellow fever, dengue, chikungunya, and Zika fevers, a strain of transgenically modified Aedes aegypti mosquitoes containing a dominant lethal gene has been developed by a commercial company, Oxitec Ltd. If lethality is complete, releasing this strain should only reduce population size and not affect the genetics of the target populations. Approximately 450 thousand males of this strain were released each week for 27 months in Jacobina, Bahia, Brazil. We genotyped the release strain and the target Jacobina population before releases began for >21,000 single nucleotide polymorphisms (SNPs). Genetic sampling from the target population six, 12, and 27–30 months after releases commenced provides clear evidence that portions of the transgenic strain genome have been incorporated into the target population. Evidently, rare viable hybrid offspring between the release strain and the Jacobina population are sufficiently robust to be able to reproduce in nature. The release strain was developed using a strain originally from Cuba, then outcrossed to a Mexican population. Thus, Jacobina Ae. aegypti are now a mix of three populations. It is unclear how this may affect disease transmission or affect other efforts to control these dangerous vectors. These results highlight the importance of having in place a genetic monitoring program during such releases to detect un-anticipated outcomes.
Hookworm infection causes anemia, malnutrition, and growth delay, especially in children living in sub-Saharan Africa. The World Health Organization recommends periodic mass drug administration (MDA) of anthelminthics to school-age children (SAC) as a means of reducing morbidity. Recently, questions have been raised about the effectiveness of MDA as a global control strategy for hookworms and other soil-transmitted helminths (STHs). Genomic DNA was extracted from Necator americanus hookworm eggs isolated from SAC enrolled in a cross-sectional study of STH epidemiology and deworming response in Kintampo North Municipality, Ghana. A polymerase chain reaction (PCR) assay was then used to identify single-nucleotide polymorphisms (SNPs) associated with benzimidazole resistance within the N. americanus β-tubulin gene. Both F167Y and F200Y resistance-associated SNPs were detected in hookworm samples from infected study subjects. Furthermore, the ratios of resistant to wild-type SNP at these two loci were increased in posttreatment samples from subjects who were not cured by albendazole, suggesting that deworming drug exposure may enrich resistance-associated mutations. A previously unreported association between F200Y and a third resistance-associated SNP, E198A, was identified by sequencing of F200Y amplicons. These data confirm that markers of benzimidazole resistance are circulating among hookworms in central Ghana, with unknown potential to impact the effectiveness and sustainability of chemotherapeutic approaches to disease transmission and control.
This paper presents an investigation into the development of an intelligent mobile-enabled expert system to perform an automatic detection of tuberculosis (TB) disease in real-time. One third of the global population are infected with the TB bacterium, and the prevailing diagnosis methods are either resource-intensive or time consuming. Thus, a reliable and easy-to-use diagnosis system has become essential to make the world TB free by 2030, as envisioned by the World Health Organisation. In this work, the challenges in implementing an efficient image processing platform is presented to extract the images from plasmonic ELISAs for TB antigen-specific antibodies and analyse their features. The supervised machine learning techniques are utilised to attain binary classification from eighteen lower-order colour moments. The proposed system is trained off-line, followed by testing and validation using a separate set of images in real-time. Using an ensemble classifier, Random Forest, we demonstrated 98.4% accuracy in TB antigen-specific antibody detection on the mobile platform. Unlike the existing systems, the proposed intelligent system with real time processing capabilities and data portability can provide the prediction without any opto-mechanical attachment, which will undergo a clinical test in the next phase. (C) 2018 The Authors. Published by Elsevier Ltd.
Giant tortoises are among the longest-lived vertebrate animals and, as such, provide an excellent model to study traits like longevity and age-related diseases. However, genomic and molecular evolutionary information on giant tortoises is scarce. Here, we describe a global analysis of the genomes of Lonesome George-the iconic last member of Chelonoidis abingdonii-and the Aldabra giant tortoise (Aldabrachelys gigantea). Comparison of these genomes with those of related species, using both unsupervised and supervised analyses, led us to detect lineage-specific variants affecting DNA repair genes, inflammatory mediators and genes related to cancer development. Our study also hints at specific evolutionary strategies linked to increased lifespan, and expands our understanding of the genomic determinants of ageing. These new genome sequences also provide important resources to help the efforts for restoration of giant tortoise populations.
Female Aedes aegypti mosquitoes infect more than 400 million people each year with dangerous viral pathogens including dengue, yellow fever, Zika and chikungunya. Progress in understanding the biology of mosquitoes and developing the tools to fight them has been slowed by the lack of a high-quality genome assembly. Here we combine diverse technologies to produce the markedly improved, fully re-annotated AaegL5 genome assembly, and demonstrate how it accelerates mosquito science. We anchored physical and cytogenetic maps, doubled the number of known chemosensory ionotropic receptors that guide mosquitoes to human hosts and egg-laying sites, provided further insight into the size and composition of the sex-determining M locus, and revealed copy-number variation among glutathione S -transferase genes that are important for insecticide resistance. Using high-resolution quantitative trait locus and population genomic analyses, we mapped new candidates for dengue vector competence and insecticide resistance. AaegL5 will catalyse new biological insights and intervention strategies to fight this deadly disease vector.
AbstractAedes aegypti, the major vector of dengue, yellow fever, chikungunya, and Zika viruses, remains of great medical and public health concern. There is little doubt that the ancestral home of the species is Africa. This mosquito invaded the New World 400‐500 years ago and later, Asia. However, little is known about the genetic structure and history of Ae. aegypti across Africa, as well as the possible origin(s) of the New World invasion. Here, we use ~17,000 genome‐wide single nucleotide polymorphisms (SNPs) to characterize a heretofore undocumented complex picture of this mosquito across its ancestral range in Africa. We find signatures of human‐assisted migrations, connectivity across long distances in sylvan populations, and of local admixture between domestic and sylvan populations. Finally, through a phylogenetic analysis combined with the genetic structure analyses, we suggest West Africa and especially Angola as the source of the New World's invasion, a scenario that fits well with the historic record of 16th‐century slave trade between Africa and Americas.
IMPORTANCE Early esophagogastric cancer (OGC) stage presents with nonspecific symptoms. OBJECTIVE The aim of this study was to determine the accuracy of a breath test for the diagnosis of OGC in a multicenter validation study. DESIGN, SETTING, AND PARTICIPANTS Patient recruitment for this diagnostic validation study was conducted at 3 London hospital sites, with breath samples returned to a central laboratory for selected ion flow tube mass spectrometry (SIFT-MS) analysis. Based on a 1:1 cancer: control ratio, and maintaining a sensitivity and specificity of 80%, the sample size required was 325 patients. All patients with cancer were on a curative treatment pathway, and patients were recruited consecutively. Among the 335 patients included; 172 were in the control group and 163 had OGC. INTERVENTIONS Breath samples were collected using secure 500-mL steel breath bags and analyzed by SIFT-MS. Quality assurance measures included sampling room air, training all researchers in breath sampling, regular instrument calibration, and unambiguous volatile organic compounds (VOCs) identification by gas chromatography mass spectrometry. MAIN OUTCOMES AND MEASURES The risk of cancerwas identified based on a previously generated 5-VOCs model and compared with histopathology-proven diagnosis. RESULTS Patients in the OGC group were older (median [IQR] age 68 [60-75] vs 55 [41-69] years) and had a greater proportion of men (134 [82.2%]) vs women (81 [47.4%]) compared with the control group. Of the 163 patients with OGC, 123 (69%) had tumor stage T3/4, and 106 (65%) had nodal metastasis on clinical staging. The predictive probabilities generated by this 5-VOCs diagnostic model were used to generate a receiver operator characteristic curve, with good diagnostic accuracy, area under the curve of 0.85. This translated to a sensitivity of 80% and specificity of 81% for the diagnosis of OGC. CONCLUSIONS AND RELEVANCE This study shows the potential of breath analysis in noninvasive diagnosis of OGC in the clinical setting. The next step is to establish the diagnostic accuracy of the test among the intended population in primary care where the test will be applied.
Female Aedes aegypti mosquitoes infect hundreds of millions of people each year with dangerous viral pathogens including dengue, yellow fever, Zika, and chikungunya. Progress in understanding the biology of this insect, and developing tools to fight it, has been slowed by the lack of a high-quality genome assembly. Here we combine diverse genome technologies to produce AaegL5, a dramatically improved and annotated assembly, and demonstrate how it accelerates mosquito science and control. We anchored the physical and cytogenetic maps, resolved the size and composition of the elusive sex-determining “M locus”, significantly increased the known members of the glutathione-S-transferase genes important for insecticide resistance, and doubled the number of chemosensory ionotropic receptors that guide mosquitoes to human hosts and egg-laying sites. Using high-resolution QTL and population genomic analyses, we mapped new candidates for dengue vector competence and insecticide resistance. We predict that AaegL5 will catalyse new biological insights and intervention strategies to fight this deadly arboviral vector.
BACKGROUND:Aedes aegypti, commonly known as "the yellow fever mosquito", is of great medical concern today primarily as the major vector of dengue, chikungunya and Zika viruses, although yellow fever remains a serious health concern in some regions. The history of Ae. aegypti in Brazil is of particular interest because the country was subjected to a well-documented eradication program during 1940s-1950s. After cessation of the campaign, the mosquito quickly re-established in the early 1970s with several dengue outbreaks reported during the last 30 years. Brazil can be considered the country suffering the most from the yellow fever mosquito, given the high number of dengue, chikungunya and Zika cases reported in the country, after having once been declared "free of Ae. aegypti". METHODOLOGY/PRINCIPAL FINDINGS:We used 12 microsatellite markers to infer the genetic structure of Brazilian Ae. aegypti populations, genetic variability, genetic affinities with neighboring geographic areas, and the timing of their arrival and spread. This enabled us to reconstruct their recent history and evaluate whether the reappearance in Brazil was the result of re-invasion from neighboring non-eradicated areas or re-emergence from local refugia surviving the eradication program. Our results indicate a genetic break separating the northern and southern Brazilian Ae. aegypti populations, with further genetic differentiation within each cluster, especially in southern Brazil. CONCLUSIONS/SIGNIFICANCE:Based on our results, re-invasions from non-eradicated regions are the most likely scenario for the reappearance of Ae. aegypti in Brazil. While populations in the northern cluster are likely to have descended from Venezuela populations as early as the 1970s, southern populations seem to have derived more recently from northern Brazilian areas. Possible entry points are also revealed within both southern and northern clusters that could inform strategies to control and monitor this important arbovirus vector.
The effective population size (Ne) is a fundamental parameter in population genetics that determines the relative strength of selection and random genetic drift, the effect of migration, levels of inbreeding, and linkage disequilibrium. In many cases where it has been estimated in animals, Ne is on the order of 10%–20% of the census size. In this study, we use 12 microsatellite markers and 14,888 single nucleotide polymorphisms (SNPs) to empirically estimate Ne in Aedes aegypti, the major vector of yellow fever, dengue, chikungunya, and Zika viruses. We used the method of temporal sampling to estimate Ne on a global dataset made up of 46 samples of Ae. aegypti that included multiple time points from 17 widely distributed geographic localities. Our Ne estimates for Ae. aegypti fell within a broad range (~25–3,000) and averaged between 400 and 600 across all localities and time points sampled. Adult census size (Nc) estimates for this species range between one and five thousand, so the Ne/Nc ratio is about the same as for most animals. These Ne values are lower than estimates available for other insects and have important implications for the design of genetic control strategies to reduce the impact of this species of mosquito on human health.
BACKGROUND:Antimicrobial Resistance is threatening our ability to treat common infectious diseases and overuse of antimicrobials to treat human infections in hospitals is accelerating this process. Clinical Decision Support Systems (CDSSs) have been proven to enhance quality of care by promoting change in prescription practices through antimicrobial selection advice. However, bypassing an initial assessment to determine the existence of an underlying disease that justifies the need of antimicrobial therapy might lead to indiscriminate and often unnecessary prescriptions.METHODS:From pathology laboratory tests, six biochemical markers were selected and combined with microbiology outcomes from susceptibility tests to create a unique dataset with over one and a half million daily profiles to perform infection risk inference. Outliers were discarded using the inter-quartile range rule and several sampling techniques were studied to tackle the class imbalance problem. The first phase selects the most effective and robust model during training using ten-fold stratified cross-validation. The second phase evaluates the final model after isotonic calibration in scenarios with missing inputs and imbalanced class distributions.RESULTS:More than 50% of infected profiles have daily requested laboratory tests for the six biochemical markers with very promising infection inference results: area under the receiver operating characteristic curve (0.80-0.83), sensitivity (0.64-0.75) and specificity (0.92-0.97). Standardization consistently outperforms normalization and sensitivity is enhanced by using the SMOTE sampling technique. Furthermore, models operated without noticeable loss in performance if at least four biomarkers were available.CONCLUSION:The selected biomarkers comprise enough information to perform infection risk inference with a high degree of confidence even in the presence of incomplete and imbalanced data. Since they are commonly available in hospitals, Clinical Decision Support Systems could benefit from these findings to assist clinicians in deciding whether or not to initiate antimicrobial therapy to improve prescription practices.