Effective management of imaging data is critical to ensuring reproducibility, accessibility, and long-term value of image-based biological research. However, good data management remains a major challenge across imaging facilities, with lack of standardised practices and limited training opportunities. Additionally, there is a need to upskill Research Technical Professionals (RTPs) in areas such as image data stewardship, as well as opening new career pathways. Here, we report on an approach to identify and design a training curriculum for imaging data stewardship in the UK bioimaging community. An initial survey captured current practices, challenges, and training requirements. These findings were presented at a community event, where structured feedback activities further identified pain points across the imaging data life cycle, and prioritised training needs. These consultations resulted in a three-module training framework covering data management and metadata, data storage and sharing, and public repository submission. The iterative approach provided technical guidance for curriculum design and community input on delivery models, target audiences, and incentives for engagement. This work highlights the need for training resources that address critical skill gaps, increase RTP visibility and recognition, and improve the quality of image data management.
Motivation In recent years, public image resources have emerged, but finding quality data efficiently remains a challenge, therefore limiting reuse.Results IDR searcher is an open-source search engine designed to facilitate the exploration of datasets hosted in public bioimaging resources. The application offers a fast, efficient, cost-effective solution for discovering datasets and has the potential to address current disparities in finding quality datasets for exploratory research and can be combined with metadata visualization tools to enhance usability for the scientific community.Availability IDR searcher is deployed using Ansible playbooks and released under the GPL v2 license. The source code associated with this manuscript is available at https://doi.org/10.5281/zenodo.20641515.
We introduce bia-binder (BioImage Archive Binder), an open-source, cloud-architectured, and web-based coding environment tailored to bioimage analysis that is freely accessible to all researchers. The service generates easy-to-use Jupyter Notebook coding environments hosted on EMBL-EBI's Embassy Cloud, which provides significant computational resources. The bia-binder architecture is free, open-source and publicly available for deployment. It features fast and direct access to images in the BioImage Archive, the Image Data Resource, and the BioStudies databases. We believe that this service can play a role in mitigating the current inequalities in access to scientific resources across academia. As bia-binder produces permanent links to compiled coding environments, we foresee the service to become widely-used within the community and enable exploratory research. bia-binder is built and deployed using helmsman and helm and released under the MIT licence. It can be accessed at binder.bioimagearchive.org and runs on any standard web browser.
We introduce the IDR searcher, an open-source search engine designed to ease the exploration of datasets hosted in public bioimaging resources. The application offers a fast, efficient, cost-effective solution for datasets discovery and has the potential to address current disparities in finding quality datasets for exploratory research and to provide the foundation for cross-resource search.
Background COVID-19 shifted Indonesian education to remote learning. The Gadjah Mada University (UGM) Faculty of Medicine Public Health and Nursing (UGM FMPHN) struggled with studying tumor images remotely due to resource shortages. In parallel, the University of Dundee's (UoD) OME team created OMERO for image management, inspiring UGM to use OMERO to develop GamaPath for web-based image viewing. This report details its implementation, effectiveness, and potential expansion to practical sessions and workshops in Indonesia. Methods Teaching slides were scanned in Indonesia and imported into OMERO at UoD. UGM students, residents, and clinicians used GamaPath for anatomical pathology training for several sessions between 2022 and 2024. GamaPath was also used for national continuing education sessions for pathologists held through 2023. Experiences and evaluations were collected via online surveys and results were assessed by modified MARuL scores. Results The UoD-UGM collaboration produced an application satisfying the initial need for remote teaching during COVID lockdown. However, after returning to in-person teaching, access to interactive training materials online was considered to be an essential part of effective instruction by students and faculty. The GamaPath application provided flexible access to interactive materials that enhanced the educational experience. Of 256 survey respondents, mean modified MARuL score for GamaPath web app was 40.92 (SD = 10.73) with a median of 41 (IQR = 34-50). Among the 107 pathologists who used GamaPath for national continuing education, the majority gave it the highest possible rating for functionality, ease of use, and overall experience. Conclusions UGMs FMPHN collaborated with UoD to create a digital pathology platform based on OMERO, improving image analysis and feedback for practical sessions and workshops using GamaPath.
A growing community is constructing a next-generation file format (NGFF) for bioimaging to overcome problems of scalability and heterogeneity. Organized by the Open Microscopy Environment (OME), individuals and institutes across diverse modalities facing these problems have designed a format specification process (OME-NGFF) to address these needs. This paper brings together a wide range of those community members to describe the cloud-optimized format itself -- OME-Zarr -- along with tools and data resources available today to increase FAIR access and remove barriers in the scientific process. The current momentum offers an opportunity to unify a key component of the bioimaging domain -- the file format that underlies so many personal, institutional, and global data management and analysis tasks.
The main goals and challenges for the life science communities in the Open Science framework are to increase reuse and sustainability of data resources, software tools, and workflows, especially in large-scale data-driven research and computational analyses. Here, we present key findings, procedures, effective measures and recommendations for generating and establishing sustainable life science resources based on the collaborative, cross-disciplinary work done within the EOSC-Life (European Open Science Cloud for Life Sciences) consortium. Bringing together 13 European life science research infrastructures, it has laid the foundation for an open, digital space to support biological and medical research. Using lessons learned from 27 selected projects, we describe the organisational, technical, financial and legal/ethical challenges that represent the main barriers to sustainability in the life sciences. We show how EOSC-Life provides a model for sustainable data management according to FAIR (findability, accessibility, interoperability, and reusability) principles, including solutions for sensitive- and industry-related resources, by means of cross-disciplinary training and best practices sharing. Finally, we illustrate how data harmonisation and collaborative work facilitate interoperability of tools, data, solutions and lead to a better understanding of concepts, semantics and functionalities in the life sciences.
SummaryBiological imaging is one of the most innovative fields in the modern biological sciences. New imaging modalities, probes, and analysis tools appear every few months and often prove decisive for enabling new directions in scientific discovery. One feature of this dynamic field is the need to capture new types of data and data structures. While there is a strong drive to make scientific data Findable, Accessible, Interoperable and Reproducible (FAIR1), the rapid rate of innovation in imaging impedes the unification and adoption of standardized data formats. Despite this, the opportunities for sharing and integrating bioimaging data and, in particular, linking these data to other “omics” datasets have never been greater. Therefore, to every extent possible, increasing “FAIRness” of bioimaging data is critical for maximizing scientific value, as well as for promoting openness and integrity.In the absence of a common, FAIR format, two approaches have emerged to provide access to bioimaging data: translation and conversion. On-the-fly translation produces a transient representation of bioimage metadata and binary data but must be repeated on each use. In contrast, conversion produces a permanent copy of the data, ideally in an open format that makes the data more accessible and improves performance and parallelization in reads and writes. Both approaches have been implemented successfully in the bioimaging community but both have limitations. At cloud-scale, those shortcomings limit scientific analysis and the sharing of results. We introduce here next-generation file formats (NGFF) as a solution to these challenges.
The rapid pace of innovation in biological imaging and the diversity of its applications have prevented the establishment of a community-agreed standardized data format. We propose that complementing established open formats such as OME-TIFF and HDF5 with a next-generation file format such as Zarr will satisfy the majority of use cases in bioimaging. Critically, a common metadata format used in all these vessels can deliver truly findable, accessible, interoperable and reusable bioimaging data.
Faced with the need to support a growing number of whole slide imaging (WSI) file formats, our team has extended a long-standing community file format (OME-TIFF) for use in digital pathology. The format makes use of the core TIFF specification to store multi-resolution (or “pyramidal”) representations of a single slide in a flexible, performant manner. Here we describe the structure of this format, its performance characteristics, as well as an open-source library support for reading and writing pyramidal OME-TIFFs.
Digital imaging is now used throughout biological and biomedical research to measure the architecture, composition and dynamics of cells, tissues and organisms. The many different imaging technologies create data in many different formats. This diversity of data formats arises because the creation