Aquatic ecosystems are vital in regulating climate and providing resources, but they face threats from global change and local stressors. Understanding their dynamics is crucial for sustainable use and conservation. The iMagine AI Platform offers a suite of AI-powered image analysis tools for researchers in aquatic sciences, facilitating a better understanding of scientific phenomena and applying AI and ML for processing image data. The platform supports the entire machine learning cycle, from model development to deployment, leveraging data from underwater platforms, webcams, microscopes, drones, and satellites, and utilising distributed resources across Europe. With a serverless architecture and DevOps approach, it enables easy sharing and deployment of AI models. Four providers within the pan-European EGI federation power the platform, offering substantial computational resources for image processing. Five use cases focus on image analytics services, which will be available to external researchers through Virtual Access. Additionally, three new use cases are developing AI-based image processing services, and two external use cases are kickstarting through recent Open Calls. The iMagine Competence Centre aids use case teams in model development and deployment, resulting in various models hosted on the iMagine AI Platform, including third-party models like YoloV8. Operational best practices derived from the platform providers and use case developers cover data management, quality control, integration, and FAIRness. These best practices aim to harmonise approaches across Research Infrastructures and will be disseminated through various channels, benefitting the broader European and international scientific communities.
As the field of machine learning advances, managing and monitoring intelligent models in production, also known as machine learning operations (MLOps), has become essential. Organizations are increasingly adopting artificial intelligence as a strategic tool, thus increasing the need for reliable, and scalable MLOps platforms. Consequently, every aspect of the machine learning life cycle, from workflow orchestration to performance monitoring, presents both challenges and opportunities that require sophisticated, flexible, and scalable technological solutions. This research addresses this demand by providing a comprehensive assessment framework of MLOps platforms highlighting the key features necessary for a robust MLOps solution. The paper examines 16 MLOps tools widely used, which revolve around capabilities within AI infrastructure management, including but not limited to experiment tracking, model deployment, and model inference. Our three-step evaluation framework starts with a feature analysis of the MLOps platforms, then GitHub stars growth assessment for adoption and prominence, and finally, a weighted scoring method to single out the most influential platforms. From this process, we derive valuable insights into the essential components of effective MLOps systems and provide a decision-making flowchart that simplifies platform selection. This framework provides hands-on guidance for organizations looking to initiate or enhance their MLOps strategies, whether they require an end-end solutions or specialized tools.
The increasing generation of data in different areas of life, such as the environment, highlights the need to explore new techniques for processing and exploiting data for useful purposes. In this context, artificial intelligence techniques, especially through deep learning models, are key tools to be used on the large amount of data that can be obtained, for example, from weather radars. In many cases, the information collected by these radars is not open, or belongs to different institutions, thus needing to deal with the distributed nature of this data. In this work, the applicability of a personalized federated learning architecture, which has been called adapFL, on distributed weather radar images is addressed. To this end, given a single available radar covering 400 km in diameter, the captured images are divided in such a way that they are disjointly distributed into four different federated clients. The results obtained with adapFL are analyzed in each zone, as well as in a central area covering part of the surface of each of the previously distributed areas. The ultimate goal of this work is to study the generalization capability of this type of learning technique for its extrapolation to use cases in which a representative number of radars is available, whose data can not be centralized due to technical, legal or administrative concerns. The results of this preliminary study indicate that the performance obtained in each zone with the adapFL approach allows improving the results of the federated learning approach, the individual deep learning models and the classical Continuity Tracking Radar Echoes by Correlation approach.
Machine learning is one of the most widely used technologies in the field of Artificial Intelligence. As machine learning applications become increasingly ubiquitous, concerns about data privacy and security have also grown. The work in this paper presents a broad theoretical landscape concerning the evolution of machine learning and deep learning from centralized to distributed learning, first in relation to privacy-preserving machine learning and secondly in the area of privacy-enhancing technologies. It provides a comprehensive landscape of the synergy between distributed machine learning and privacy-enhancing technologies, with federated learning being one of the most prominent architectures. Various distributed learning approaches to privacy-aware techniques are structured in a review, followed by an in-depth description of relevant frameworks and libraries, more particularly in the context of federated learning. The paper also highlights the need for data protection and privacy addressed from different approaches, key findings in the field concerning AI applications, and advances in the development of related tools and techniques.
Open Science is a paradigm in which scientific data, procedures, tools and results are shared transparently and reused by society as a whole. The initiative known as the European Open Science Cloud (EOSC) is an effort in Europe to provide an open, trusted, virtual and federated computing environment to execute scientific applications, and to store, share and re-use research data across borders and scientific disciplines. Additionally, scientific services are becoming increasingly data-intensive, not only in terms of computationally intensive tasks but also in terms of storage resources. Computing paradigms such as High Performance Computing (HPC) and Cloud Computing are applied to e-science applications to meet these demands. However, adapting applications and services to these paradigms is not a trivial task, commonly requiring a deep knowledge of the underlying technologies, which often constitutes a barrier for its uptake by scientists in general. In this context, EOSC-SYNERGY, a collaborative project involving more than 20 institutions from eight European countries pooling their knowledge and experience to enhance EOSC's capabilities and capacities, aims to bring EOSC closer to the scientific communities. This article provides a summary analysis of the adaptations made in the ten thematic services of EOSC-SYNERGY to embrace this paradigm. These services are grouped into four categories: Earth Observation, Environment, Biomedicine, and Astrophysics. The analysis will lead to the identification of commonalities, best practices and common requirements, regardless of the thematic area of the service. Experience gained from the thematic services could be transferred to new services for the adoption of the EOSC ecosystem framework.
<p>The O3as service is a tool designed to support the assessment of atmospheric ozone levels and trends. It was developed as one of the thematic services of the EOSC-Synergy project. It allows for the analysis of large datasets from chemistry-climate models and presents the information in a user-friendly format for a broad range of users, including scientists, pupils, and interested citizens. The service utilizes a unified approach to process the data, employs CF conventions for homogenization, and generates figures that can be published or downloaded as csv files. It was developed as part of the EOSC-Synergy project, and it runs on a cloud-based, containerized architecture orchestrated by Kubernetes and HPC resources, and uses the Large Scale Data Facility (LSDF) at the KIT for data storage. The service is developed with best software practices, including quality assurance, continuous integration and delivery, and compliance with the FAIR principles.&#160;</p> <p>This presentation will focus in particular on the architecture and functionality of the O3as service, with an example demonstration of its usage.</p>
Modern digital scientific workflows - often implying Big Data challenges - require data infrastructures and innovative data science methods across disciplines and technologies. Diverse activities within and outside HGF deal with these challenges, on all levels. The series of Data Science Symposia fosters knowledge exchange and collaboration in the Earth and Environment research community. We invited contributions to the overarching topics of data management, data science and data infrastructures. The series of Data Science Symposia is a joint initiative by the three Helmholtz Centers HZG, AWI and GEOMAR Organization: Hela Mehrtens and Daniela Henkel (GEOMAR)
The European Open Science Cloud-Synergy (EOSC-Synergy) project delivers services that serve to expand the use of EOSC. One of these services, O3as, is being developed for scientists using chemistry-climate models to determine time series and eventually ozone trends for potential use in the quadrennial Global Assessment of Ozone Depletion, which will be published in 2022. A unified approach from a service like ours, which analyses results from a large number of different climate models, helps to harmonise the calculation of ozone trends efficiently and consistently. With O3as, publication-quality figures can be reproduced quickly and in a coherent way. This is done via a web application where users configure their queries to perform simple analyses. These queries are passed to the O3as service via an O3as REST API call. There, the O3as service processes the query and accesses the reduced dataset. To create a reduced dataset, regular tasks are executed on a high performance computer (HPC) to copy the primary data and perform data preparation (e.g. data reduction, standardisation and parameter unification). O3as uses EGI check-in (OIDC) to identify users and grant access to certain functionalities of the service, udocker (a tool to run Docker containers in multi-user space without root privileges) to perform data reduction in the HPC environment, and the Universitat Politècnica de València (UPV) Infrastructure Manager to provision service resources (Kubernetes).
EOSC-SYNERGY receives funding from the European Union’s Horizon 2020 research and innovation programme under grant agreement No. 857647.
Ever growing interest and usage of deep learning rises a question on the performance of various infrastructures suitable for training of neural networks. We present here our approach and first results of tests performed with TensorFlow Benchmarks which use best practices for multi-GPU and distributed training. We pack the Benchmarks in Docker containers and execute them by means of uDocker and Singularity container tools on a single machine and in the HPC environment. The Benchmarks comprise a number of convolutional neural network models run across synthetic data and e.g. the ImageNet dataset. For the same Nvidia K80 GPU card we achieve the same performance in terms of processed images per second and similar scalability between 1-2-4 GPUs as presented by the TensorFlow developers. We therefore do not obtain statistically significant overhead due to the usage of containers in the multi-GPU case, and the approach of using TF Benchmarks in a Docker container can be applied across various systems.
The Edelweiss experiment uses Ge-bolometers with an improved background rejection (interleaved electrode design) to search for WIMP dark matter. The setup is located in the underground laboratory, Laboratoire Souterrain de Modane (LSM, France). In 20092010 the collaboration successfully operated ten 400-g bolometers together with an active muon veto shielding. Published analysis of this measurement campaign was optimized for WIMP masses above 50 GeV. Recently, the analysis was extended to the low-mass WIMP region using a quality subset of the 2009-2010 data setting new limits on the spinindependent WIMP-nucleon scattering cross-section. We present the low-mass WIMP analysis, background investigations and the latest measurements with a subset of the forty 800-g detectors that will be installed for the Edelweiss-III. Ongoing installation works of the Edelweiss-III setup and further plans for a next generation experiment, EURECA, are discussed.
The first high-statistics and high-resolution data set for the integrated recoil-ion energy spectrum following the \( \beta^+\) decay of 35Ar has been collected with the WITCH retardation spectrometer located at CERN-ISOLDE. Over 25 million recoil-ion events were recorded on a large-area multichannel plate (MCP) detector with a time-stamp precision of 2ns and position resolution of 0.1mm due to the newly upgraded data acquisition based on the LPC Caen FASTER protocol. The number of recoil ions was measured for more than 15 different settings of the retardation potential, complemented by dedicated background and half-life measurements. Previously unidentified systematic effects, including an energy-dependent efficiency of the main MCP and a radiation-induced time-dependent background, have been identified and incorporated into the analysis. However, further understanding and treatment of the radiation-induced background requires additional dedicated measurements and remains the current limiting factor in extracting a beta-neutrino angular correlation coefficient for 35Ar decay using the WITCH spectrometer.
The EDELWEISS experiment, located in the underground laboratory LSM (France), is one of the leading experiments using cryogenic germanium (Ge) detectors for a direct search for dark matter. For the EDELWEISS-III phase, a new scalable data acquisition (DAQ) system was designed and built, based on the `IPE4 DAQ system', which has already been used for several experiments in astroparticle physics.
A dedicated analysis of the muon-induced background in the EDELWEISS dark matter search has been performed on a data set acquired in 2009 and 2010. The total muon flux underground in the Laboratoire Souterrain de Modane (LSM) was measured to be Φμ=(5.4±0.2-0.9+0.5) muons/m2/d. The modular design of the μ-veto system allows the reconstruction of the muon trajectory and hence the determination of the angular dependent muon flux in LSM. The results are in good agreement with both MC simulations and earlier measurements. Synchronization of the μ-veto system with the phonon and ionization signals of the Ge detector array allowed identification of muon-induced events. Rates for all muon-induced events Γμ=(0.172±0.012)evts/(kgd) and of WIMP-like events Γμ–n=0.008-0.004+0.005evts/(kgd) were extracted. After vetoing, the remaining rate of accepted muon-induced neutrons in the EDELWEISS-II dark matter search was determined to be Γirredμ–n<6·10-4evts/(kgd) at 90% C.L. Based on these results, the muon-induced background expectation for an anticipated exposure of 3000 kg d for EDELWEISS-III is N3000kgdμ–n<0.6 events.
Due to a very low event rate expected in direct dark matter search experiments, a good understanding of every background component is crucial. Muon-induced neutrons constitute a prominent background, since neutrons lead to nuclear recoils and thus can mimic a potential dark matter signal. EDELWEISS is a Ge-bolometer experiment searching for WIMP dark matter. It is located in the Laboratoire Souterrain de Modane (LSM, France). We have measured muon-induced neutrons by means of a neutron counter based on Gd-loaded liquid scintillator. Studies of muon-induced neutrons are presented and include development of the appropriate MC model based on Geant4 and analysis of a 1000-days measurement campaign in LSM. We find a good agreement between measured rates of muon-induced neutrons and those predicted by the developed model with full event topology. The impact of the neutron background on current EDELWEISS data-taking as well as for next generation experiments such as EURECA is briefly discussed.