
Reproducibility challenges scientific reporting, including metagenomics, where increasingly complex bioinformatics pipelines hinder transparency, comparability, and customization for life science students globally. To address this demanding task, we built an open, interactive and web-based tutorial that guides scholars with basic command-line skills through the detailed development of a validated and reproducible Nextflow metagenomics classification pipeline. As important features, the tutorial emphasizes simplicity, modularity, and containerization, which empowers users with both conceptual understanding and practical implementation skills. Noteworthy, this tutorial provides all the required files, databases, dependencies, software and environment for users to run it without the need of local installation or computational adaptations elsewhere. Finally, by offering a fully reproducible pipeline with a step-by-step developing tutorial, this work aims to lower technical barriers in microbiome bioinformatics and promote best practices in metagenomics data analysis. TaxoFlow is freely available at https://taxoflow.work/.
Abstract This data paper presents a curated, georeferenced dataset of the frequencies of the two main target site mutations (V1016G and F1534C) associated with resistance to pyrethroid insecticides in Aedes albopictus in Italy. Populations were collected in 102 out 107 Italian provinces between 2023 and 2025. Specimens were sampled by members of the Mosquito Insecticide Resistance Italian Network (MosqIRIT) as part of RN2 activities within the INF-ACT project. Genotyping was performed on 3,517 individuals by specific allele-specific PCR assays. Each record includes metadata on sampling site, administrative location, developmental stage, collection method, and mutation-specific genotype frequencies. To support spatial analysis modelling effort, the dataset integrates geographic, eco-climatic, and demographic data. This resource will support mosquito control programs, pyrethroid resistance monitoring and managing, as well as ecological modelling, and is compliant with the FAIR data program.
We present ChestPathCT5-S100, an open dataset of 87 real-world chest CT examinations spanning five common thoracic pathologies: rib fracture, pleural effusion, lung mass, pulmonary embolism, and pneumothorax. The dataset was assembled from a retrospective single-centre cohort over a ten-year period, intentionally preserving acquisition heterogeneity and concomitant findings representative of routine clinical practice. Cases include both contrast-enhanced (arterial and venous phase) and non-contrast examinations, drawn from emergency, oncologic, and trauma settings. All imaging volumes are provided in NIfTI format. Technical validation by two radiologists confirmed correct pathology category assignment, image integrity, and unambiguous visibility of the dominant pathology for each case. A binary co-occurrence matrix of concomitant findings is provided to support multi-label research designs. ChestPathCT5-S100 is publicly available on Zenodo under a CC0 1.0 license, permitting unrestricted use and redistribution. The dataset supports classification, detection, weakly supervised learning, and multi-task learning paradigms.
Abstract Widow spiders of the genus Latrodectus are important animals for biomedical, pest and conservation research. Here, we present the assembled genomes of two closely related Latrodectus species: the Australian L. hasselti and the New Zealand endemic L. katipo. The genome of L. katipo consists of 13 scaffolds likely corresponding to chromosomes (90% of the total length) and 1267 short scaffolds (10%). It has a total length of 1.5 Gbp and BUSCO of 94.9%. The genome of L. hasselti consists of 379 scaffolds and has a total length of 1.7 Gbp and a BUSCO score of 95.4%. The repeat content is very similar in both genomes with a total proportion of 37.2% for L. katipo and 39.9% for L. hasselti. Genome annotation predicted 12706 and 15111 genes for L. katipo and L. hasselti respectively. An ortholog analysis shows large overlap between orthogroups suggesting either duplication events in L. hasselti or loss of genes in L. katipo.
We present AortaSeg-60, an open dataset of 60 real-world thoraco-abdominal CT-angiography scans of the aorta encompassing normal anatomy and pathological variations, designed for AI research, benchmarking, and educational purposes. The dataset is organized into six balanced categories: young normal, elderly normal, aortic aneurysms, aortic dissections, venous acquisition, and non-contrast acquisition, capturing realistic anatomical and pathological diversity. All scans are provided in NIfTI format with fully automated aortic segmentation masks generated using TotalSegmentator, without manual correction, enabling evaluation of typical algorithmic errors and testing of refinement strategies. Two radiologists performed a technical validation to ensure dataset curation and correct category assignment. AortaSeg-60 is publicly available on Zenodo under a CC0 license. By providing paired imaging and automated labels, the dataset facilitates reproducible research, algorithm development, and method comparison for vascular segmentation, while noting limitations of sample size, single-centre acquisition, and reliance on automated annotations.
High-throughput sequencing datasets frequently exhibit extreme read depth variation, biasing downstream analysis. Normalising coverage to a specific depth cap is important, yet existing tools rely on computationally expensive fetch-based or non-deterministic greedy algorithms. Here, we present a new coordinate-sorted sweep-line algorithm implemented in the open-source software rasusa that enforces a strict coverage cap at every genomic position. By utilising seeded random priority assignment, we achieve unbiased, reproducible read selection. The algorithm reduces runtimes by over 1,400-fold compared to legacy fetch-based methods—slashing processing from hours to mere seconds—and operates roughly four times faster than VariantBam. Furthermore, it requires only 8 MB of memory for long-read data. This provides a highly efficient, scalable, and reproducible solution for sequencing coverage normalisation.
This dataset arises from a multilingual survey of AI use among participants and community members in the DBCLS BioHackathon 2025 in Japan. The questionnaire, offered in English, Japanese, and Thai, asked about how often respondents use AI tools, what they use them for, obstacles they encounter, institutional support, satisfaction, and concerns. Additional items captured role, institution type, work country, and other demographics, totaling 105 responses. The dataset includes both raw anonymized responses and a cleaned, standardized English-only version suitable for quantitative analysis, along with the full questionnaire, a data dictionary for cleaned dataset, and a translation lookup table. Free-text answers were screened and redacted to remove URLs, names, and other potentially identifiable information. Together, these materials provide a community-level view of AI practice in genomics, bioinformatics, software development, and related areas, and can support work on AI adoption, policy, and methods for analyzing survey data on AI use in science.
Vector control is a cornerstone for malaria management in Sub-Saharan Africa. Understanding the distribution dynamics and ecology of major malaria vectors is important for strengthening the current control efforts of national malaria control programmes. This project monitored the spatiotemporal distribution of Anopheles mosquitoes across different ecological zones of Ghana from 2017 to 2025. Anopheles mosquitoes were sampled from twelve sites across the three ecological zones of Ghana (Coastal, Forest and Sahel Savannah zones) using human landing catches and Prokopack aspirators. Mosquitoes were subjected to morphological and molecular species identification. Sporozoite infection rates and blood meal sources of collected blood fed female mosquitoes were both assessed using PCR. A total of 47,771 Anopheline mosquitoes (An. gambiae s.l, An. funestus, An. pharoensis and An. rufipes) were collected across the three ecological zones. Anopheles gambiae s.l, and particularly An. coluzzii and An. gambiae s.s were the predominant species across the study sites and ecological zones. Sporozoite infections were higher in the forest and sahel zones compared to the coastal zone, and the overall human blood index was 40.46%. Our findings provide relevant data for improving current vector control for malaria in Ghana.
Urban malaria is an emerging challenge in sub-Saharan Africa, driven by unplanned urbanization, irrigation, and vector adaptation; yet data on urban vectors, their diversity and malaria transmission potential are limited. We assessed Anopheles gambiae s.l. abundance, species composition, and behavior in Accra, Ghana, during dry and rainy seasons of 2022 to 2024 across fifteen sites representing different socioeconomic settings. A total of 20,945 host-seeking and 1,613 resting Anopheles mosquitoes were collected. Abundance was highest in irrigation and peri-urban sites, and lowest in low socioeconomic areas. An. gambiae s.s. dominated host-seeking populations, while An. coluzzii dominated resting ones. Findings highlight irrigation and peri-urban areas as hotspots, requiring targeted surveillance and control.
Invasive species are one of the biggest drivers of species extinction. Ludwigia grandiflora subsp. hexapetala (Lgh) is widely invasive in aquatic ecosystems of Europe, North America, and Japan, and also colonizes emergent freshwater soils, but limited genomic data constrain studies of its invasiveness. Here, we report a draft genome assembly of Lgh, with a total length of 1.487 Gb, in agreement with the genome size estimated by flow cytometry, despite high fragmentation (111,219 contigs; N50 = 13.5 kb) and low sequencing depth (6.5× Illumina, 1.6× Nanopore). In addition, an analysis combining homology and expression data identified 139,095 protein-coding genes. Moreover, several indicators suggest that the observed fragmentation is largely attributable to unassembled repetitive regions. Thus, despite these limitations, this assembly represents the first genome in the Ludwigioideae subfamily and constitutes a valuable resource for gene discovery, functional genomics, phylogenetic reconstruction, and evolutionary analyses across the Onagraceae family.
We present a genome assembly from the coral species Porites harrisoni from the southern Persian/Arabian Gulf, the hottest ocean basin where corals live. The assembly is 626.7 Mb in size, spanning 1,883 contigs with a contig N50 of 807.4 kb, including a single-contig mitochondrial genome. The assembly has a BUSCO completeness of 86.3% (single = 72.5%, duplicated = 13.7%, fragmented = 1.2%, missing = 12.5%). Within the nuclear genome, 59.23% are repeats (15.89% retroelements, 10.00% DNA transposons, and 31.71% unclassified repeats). Gene annotation of the nuclear genome assembly identified 27,823 protein-coding genes. The mitogenome is 18,639 bp long, with 13 protein-coding genes, 2 tRNAs, and 2 rRNAs. The P. harrisoni genome is a valuable resource of a coral from an extreme environment, enhancing understanding of the genomic architecture underlying thermal resilience. Comparative analyses will help elucidate the evolutionary basis of heat tolerance and the adaptive capacity of coral to rapid climate change.
ABSTRACT Pacific herring (Clupea pallasii) serve as a critical trophic link between plankton and many marine species targeted by fisheries. With a broad distribution throughout the North Pacific Ocean, from the Arctic to temperate latitudes, herring hold ecological, economic, and cultural importance. Despite this importance, genomic resources for this species, such as reference genome sequences, have only recently become available. To date, only one scaffold-level reference genome, representing a specimen from the Gulf of Alaska (Vancouver; 1,379 scaffolds), has been published to NCBI. Addressing this data gap, we produced a high quality 795 Mb genome sequence organized into 26 chromosomes combining long read sequencing with short read sequencing of proximity ligation libraries. Our assembly is highly complete (BUSCO score of 97.7%) and contiguous (922 contigs, N50 = 7,338,470, L50 = 38; 26 scaffolds, N50 = 31,494,017; L50 = 12). Pacific herring from the Bering Sea are genetically differentiated from those south of the Aleutian Islands and the Alaska Peninsula, making a reference genome from the eastern Bering Sea an important addition to the Pacific herring’s genomic toolbox.
Background Understanding the genetic architecture of domestic dogs provides unique insights into the processes of domestication, breed formation, and the genetic basis of complex traits and diseases. Dog populations, characterized by their diverse morphologies and behaviors, also exhibit extensive evidence of historical and ongoing admixture. This widespread mixing, driven by both natural migration and selective breeding practices, has profoundly shaped the genomic landscape of modern dog breeds. Though global admixture has been extensively estimated in human population studies, where the number of subgroups is typically limited, there has been more limited analysis in canines, where there may be dozens of ancestral groups, or breeds. Results Here we present a procedure for estimating global admixture in dogs from whole genome sequence data using SCOPE. We created a reference population of 65 dog breeds that included 349 individuals, from which we determined breed-informative SNPs. We demonstrate that SCOPE can accurately infer breed composition in both simulated and real admixed samples, even at low sequencing depths. We also characterized the genetic similarity between our reference dog breeds and recovered previously reported relationships. Conclusion This approach allows us to identify the strength of the genetic signature of breeds and place error bounds on admixture estimates. It also provides evidence that admixture can be accurately inferred in subjects that may originate from multiple ancestral populations.
Malaria control in Ghana and sub-Saharan Africa is threatened by widespread insecticide resistance in Anopheles gambiae s.l., undermining the effectiveness of long-lasting insecticidal nets and indoor residual spraying. A longitudinal survey was conducted between 2023 and 2025 across 20 urban and suburban sites spanning the coastal savannah, forest, and Sahel savannah zones. Of the 1,008 An. gambiae s.l. sampled, An. coluzzii was the dominant species (65.1%), followed by An. gambiae s.s. (18.9%) and An. arabiensis (10.9%). WHO bioassays revealed high pyrethroid resistance (mortality rate = 20–45%). Full susceptibility to pirimiphos-methyl (mortality rate = 99–100%) and chlorfenapyr was observed at most sites, though resistance to clothianidin was observed in Obuasi, Tema, and Abossey Okai. Intensity assays confirmed strong pyrethroid resistance even at 10× diagnostic concentrations. Genotyping showed near-fixation of the kdrL995F allele and the presence of additional resistance markers, including N1570Y, V402L, I1527T, and Ace-1R G280S.
Oxygen availability is a key regulator of cellular physiology and hypoxia plays a central role driving vasculogenesis and angiogenesis during development. Although bulk transcriptomics has revealed important oxygen-regulated gene networks, such approaches cannot resolve the cellular heterogeneity and lineage dynamics characteristic of early differentiation. To address this, we generated a single-cell transcriptomic dataset from murine embryoid bodies, a widely used in vitro model of early embryonic development, cultured 8 or 10 days under hypoxic (1% O2) or normoxic (21% O2) conditions for the final 16 or 48 hours of differentiation. This resource enables detailed exploration of how oxygen availability influences lineage specification, vascular and hematopoietic development, and cellular heterogeneity during early differentiation. Beyond developmental biology, the dataset provides a valuable reference for comparative studies of hypoxia responses, benchmarking of single-cell analysis methods, and integrative investigations into oxygen signaling across diverse biological systems.
Spatial proteomics provides a spatially resolved view of protein expression and localization within cells and tissues by mapping the location and abundance of proteins. There is a need for fully-integrated end-to-end imaging workflows for spatial proteomic analysis that are flexible, reproducible, and support graphical and interactive visualizations. We present a modular and interactive spatial proteomic image analysis workflow with individual containerized steps that empowers biomedical researchers to reproducibly execute and customize complex analyses. Our workflow consists of cell segmentation, unsupervised clustering with optional batch correction, validation of clusters on the image, and cell type clustering results visualization. A form-based graphical interface can be utilized to execute and customize multi-step workflows with a single click or interactively adjust image processing steps within the workflow, apply workflows to various datasets, and modify input parameters as needed. We illustrated the functionality of our workflow using human normal tonsil and colorectal cancer tissues stained by high-plex immunohistochemistry.
Arboviral diseases such as dengue, chikungunya, Zika, and yellow fever are of increasing endemicity and public health concern in Africa. Understanding the spatial distribution and dynamics of insecticide resistance in the Aedes vector could guide effective control interventions. We conducted larval surveys and WHO adult susceptibility bioassays on emerged adults from January 2019 to December 2023 in Ghana. Bioassays revealed widespread resistance in Ae. aegypti to pyrethroids, with 33.8-88.8% mortality for deltamethrin and 65-89% for permethrin. Ae. aegypti from Paga, Takoradi, and Accra was susceptible to pirimiphos-methyl. Ae. vittatus exhibited confirmed or possible resistance to pyrethroids. Ae. albopictus was found susceptible to all insecticides tested. Genotyping of mosquitoes (n = 887) identified high allelic frequencies of the F1534C kdr mutation in the pyrethroid-resistant Ae. aegypti populations. These findings highlight widespread pyrethroid resistance in the Ghanaian Aedes populations driven primarily by target-site insensitivity, and emphasize the urgent need for evidence-based vector-management strategies.
In Africa, Culex is an important vector that transmits West Nile virus, whilst Aedes mosquitoes transmit dengue, yellow fever, chikungunya, and Zika. However, very limited data is available on their bionomics and ecology. Here, we provide data on the abundance and distribution of Culex and Aedes mosquitoes in Ghana between 2017 and 2025. We collected 39,761 Culex and 6,047 Aedes mosquitoes using various mosquito-trapping tools. Both vectors were predominantly observed outdoors. Aedes aegypti was the most dominant Aedes vector observed in Ghana. The invasive Aedes albopictus was sampled in 2023, whereas Aedes vittatus was observed in Accra. Our data provides important information to support vector surveillance, ecological risk assessments, and integrated vector-management strategies.
From May to October 2024, Cuba experienced an outbreak of Oropouche virus (OROV), an Orthobunyavirus previously restricted to the Amazon region. As no Orthobunyavirus circulation had been previously reported in Cuba, the local vector involvement was uncertain. Entomo-virological surveys were conducted in active transmission areas across three provinces. Adult insects collected with traps and aspirators were screened for OROV by real-time RT-qPCR. A total of 2,180 specimens representing six dipteran species or families were identified. Culex quinquefasciatus and Aedes aegypti occurred in all provinces, with Cx. quinquefasciatus predominating (n = 1,785), followed by Ae. aegypti (n = 285) and Ceratopogonidae (n = 49). Eleven pools containing these taxa tested positive for OROV RNA. Detection of OROV in various species suggests possible involvement of multiple vectors in the Cuban outbreak. Further studies are needed to assess vector competence and elucidate OROV transmission dynamics in the Caribbean region.
Identifying differentially expressed genes associated with genetic pathologies is crucial to understanding the biological differences between healthy and diseased states and identifying potential biomarkers and therapeutic targets. However, gene expression profiles are controlled by various mechanisms, including epigenomic changes, such as DNA methylation, histone modifications, and interfering microRNA silencing. We developed a novel Shiny application for transcriptomic and epigenomic change identification and correlation using a combination of Bioconductor and CRAN packages. The developed package, named EMImR, is a user-friendly tool with an easy-to-use graphical user interface to identify differentially expressed genes, differentially methylated genes, and differentially expressed interfering microRNA. In addition, it identifies the correlation between transcriptomic and epigenomic modifications and performs the ontology analysis of genes of interest. The developed tool could be used to study the regulatory effects of epigenetic factors. The application is publicly available in the GitHub repository (https://github.com/omicscodeathon/emimr).