Retroviruses that colonize the host germline can be passed on as inherited genetic variants. The koala (Phascolarctos cinereus) is currently experiencing germline colonization by two retroviruses, the koala retrovirus (KoRV) and phaCin-β. We analyze the integration site segregation and diversity of endogenous KoRV, phaCin-β, and the related phaCin-β-like in 111 pedigreed koalas from the San Diego Zoo Wildlife Alliance and seven European Zoos. The use of multigenerational pedigrees and the inclusion of health information for each individual koala reveal elimination of retroviruses from proto-oncogenes and the generation and spread of new germline integrations. Seven-hundred-and-fourteen integrations do not persist in the living population. For the 55 triads examined, 21 unique integrations identified in individual koalas are absent in their parents. Retroviral integrations associated with leukemia, fertility, and longevity are used to estimate genetic risk scores and develop a longevity breeding index to minimize neoplasia risk in the captive koala population.
The fused multiply-add (FMA) instruction enables the radix-2 FFT butterfly to be computed in 6 FMA operations – the proven minimum. The classical factorization by Linzer and Feig precomputes the ratio = /, which is singular when the twiddle factor is W^0 = 1 (i.e., = 0). Standard practice clamps to a small epsilon, degrading numerical precision. We observe that an alternative factorization using as the outer multiplier (precomputing ) avoids this particular singularity but introduces a new one at W^N/4. We then propose a dual-select strategy that chooses, per twiddle factor, whichever factorization yields |ratio| ≤ 1. This eliminates all singularities, requires no epsilon clamping, and bounds the precomputed ratio to unity for all twiddle factors. For N = 1024, the worst-case ratio drops from 163 (Linzer-Feig) to exactly 1.0 (dual-select), yielding a 235× tighter error bound in FP16 arithmetic over 10 FFT passes. The strategy adds zero computational overhead – only the precomputed twiddle table changes.
Whole genome sequence (WGS) data in multi-ancestry samples supports discovery of low-frequency or population-specific genetic variants associated with chronic obstructive pulmonary disease (COPD) and lung function. We performed single variant, structural variant, and gene-based analysis of pulmonary function (FEV1, FVC and FEV1/FVC) and COPD case–control status in 44,287 multi-ancestry participants from the NHLBI Trans-Omics for Precision Medicine (TOPMed) Program. We validated findings using the UK Biobank and assessed implicated genes using lung single-cell RNA-seq (scRNA-seq) data sets. Applying a genome-wide significance threshold (P < 5 × 10–9), we replicated known loci and identified novel associations near LY86, MAGI1, GRK7, and LINC02668. Colocalization with gene expression quantitative trait loci (eQTL) from the Lung Tissue Research Consortium highlighted known candidate genes including ADAM19, THSD4, C4B, and PSMA4, which were not identified through other eQTL sources. Multi-ancestry analysis improved fine-mapping resolution (e.g., HTR4 and RIN3). Gene-based analysis identified and replicated HMCN1. In human lung scRNA-seq data sets, lung epithelial cells and immune cell types showed enriched expression, while fibroblasts showed higher expression for HMCN1. CRISPR targeting HMCN1 in IMR90 demonstrated reduced expression of collagen genes. Large-scale multi-ancestry WGS analysis improves variant discovery and fine-mapping resolution for lung function and COPD and highlights biologically relevant genes and pathways.
Fetal copy-number variants (CNVs) have been associated with a broad range of phenotypes and pregnancy outcomes. Noninvasive prenatal screening using genome-wide cell-free (cf) DNA analysis offers an opportunity to detect fetal CNVs early in pregnancy. This retrospective cohort study evaluated concordance between genome-wide cfDNA screening and diagnostic test results for 276 cases with a single isolated cfDNA-identified CNV ≥7 Mb. Cases for this study were submitted by members of the Global Expanded NIPT Consortium. Eight consortium sites in seven countries contributed cases, with 83% of cases submitted from European sites. Seventy-three of 276 cases (26.5%) had no known high-risk indication for cfDNA screening. Mean and median gestational age at the time of cfDNA blood draw was 13 weeks. A deletion was identified for 124 (44.9%) cases and a duplication for 152 (55.1%) cases. Mean CNV size was 33.4 Mb (median 23.1 Mb, range 7–187 Mb). Diagnostic test results were available for 209/276 cases (75.7%). Concordance between cfDNA screening and diagnostic test results was observed for 49/209 cases (23.4%). Mean and median fetal fraction among concordant cases was 8.8% and 8%, respectively. Among 157 discordant cases, a plausible maternal biological explanation was identified for 21 cases (13.4%). Pregnancy outcome information was limited, but available for 116 (42.0%) cases. Parental testing results were available for 38 (13.8%) cases. For six of 15 concordant cases with parental results, fetal CNVs were secondary to a parental translocation or rearrangement. This study contributes to the growing evidence supporting the use of genome-wide cfDNA screening for detection of large fetal CNVs that could affect the current pregnancy and future reproductive risks as well as identify previously unknown maternal conditions.
Metatranscriptomic (MetaT) sequencing provides insights into gene expression and functional activity within microbial communities, but its utility is limited by the high abundance of ribosomal RNA (rRNA), which often accounts for ≥ 90