To investigate whether machine learning (ML) algorithms can learn racial or ethnic information from the vital signs alone. A retrospective cohort study of critically ill patients between 2014 and 2015 from the multicentre eICU-CRD critical care database involving 335 intensive care units in 208 US hospitals, containing 200 859 admissions. We extracted 10 763 critical care admissions of patients aged 18 and over, alive during the first 24 hours after admission, with recorded race or ethnicity as well as at least two measurements of heart rate, oxygen saturation, respiratory rate and blood pressure. Pairs of subgroups were matched based on age, gender, admission diagnosis and disease severity. XGBoost, Random Forest and Logistic Regression algorithms were used to predict recorded race or ethnicity based on the values of vital signs. Models derived from only four vital signs can predict patients' recorded race or ethnicity with an area under the curve (AUC) of 0.74 (±0.030) between White and Black patients, AUC of 0.74 (±0.030) between Hispanic and Black patients and AUC of 0.67 (±0.072) between Hispanic and White patients, even when controlling for known factors. There were very small, but statistically significant differences between heart rate, oxygen saturation and blood pressure, but not respiration rate and invasively measured oxygen saturation. ML algorithms can extract racial or ethnicity information from vital signs alone across diverse patient populations, even when controlling for known biases such as pulse oximetry variations and comorbidities. The model correctly classified the race or ethnicity in two out of three patients, indicating that this outcome is not random. Vital signs embed racial information that can be learnt by ML algorithms, posing a significant risk to equitable clinical decision-making. Mitigating measures might be challenging, considering the fundamental role of vital signs in clinical decision-making.
This work introduces the integration of PatchTST, a transformer-based time-series forecasting model, in the near-real-time RAN Intelligent Controller for predictive base station reconfiguration in a 5G/6G network. Additionally, an in-domain transfer learning approach is proposed, achieving 16.87% improvement in forecasting performance by leveraging only 0.15% of the dataset, while identifying a saturation point in model performance relative to the number of base stations used for transferring knowledge. Results demonstrate that PatchTST outperforms traditional machine learning baselines by 73.44%, providing highly accurate predictions. This enables effective dynamic reconfiguration of base stations and enhancing network efficiency in next-generation cellular systems.
In this letter, we study distributed federated learning (FL) in wireless powered communication networks (WPCNs). The proposed system model ensures data privacy and energy self-sustainability of wireless (e.g., sensory, sensing or data gathering) devices involved in collaborative machine learning regardless of the specific FL algorithm. We specifically aim to minimize the total training duration of the FL process by properly allocating the communication resources (i.e., durations of energy harvesting, local processing and transmission phases, and transmit powers), the computational parameters of the EH clients (i.e. CPU frequencies) and learning parameters of their FL models (i.e. local training error threshold). We derive analytical solutions for these parameters, resulting in low complexity in implementing the proposed scheme.
Split Learning (SL) is a promising Distributed Learning approach in electromyography (EMG) based prosthetic control, due to its applicability within resource-constrained environments. Other learning approaches, such as Deep Learning and Federated Learning (FL), provide suboptimal solutions, since prosthetic devices are extremely limited in terms of processing power and battery life. The viability of implementing SL in such scenarios is caused by its inherent model partitioning, with clients executing the smaller model segment. However, selecting an inadequate cut layer hinders the training process in SL systems. This paper presents an algorithm for optimal cut layer selection in terms of maximizing the convergence rate of the model. The performance evaluation demonstrates that the proposed algorithm substantially accelerates the convergence in an EMG pattern recognition task for improving prosthetic device control.
This research paper focuses on developing a complete system for daily automation testing of comprehensive web applications implemented on cloud environments, encompassing the execution of automated API tests, real-time monitoring and results visualization of the testing environments. Despite the tools for developing automated API tests, the study uses containerization tools as Docker and Kubernetes, showcasing their integration into a cohesive testing framework. Furthermore, the implementation leverages the potential of the Google Cloud Platform (GCP) to demonstrate the usage of cloud computing services, emphasizing scalability and efficiency. Additionally, the paper details the integration of monitoring tools, specifically Elasticsearch, to assess and visualize the health and performance of the underlying Kubernetes cluster. Through a comprehensive approach, encompassing a wide variety of tools, the research establishes a continuous and automated testing environment essential for cutting-edge software applications. Results showcase the successful orchestration of all the technologies, highlighting their collective impact on achieving a robust and efficient system for continuous automation testing and monitoring.
Federated learning (FL) emerges as a highly pursued approach to machine learning (ML), that enables central model training while maintaining data decentralization and privacy. The distributed computation aspect of FL makes it more appealing for scenarios with constrained bandwidth and energy, which necessitates the usage of more efficient multiple access technologies, such as non-orthogonal multiple access (NOMA), and energy harvesting (EH) technologies, such as wireless power transfer. In this paper, we thus focus on a NOMA-based wireless system with radio frequency (RF) EH for training FL models. The EH users (EHUs) exploit the harvested energy for training their respective local FL models and for transmitting the local model parameters to the base station (BS). Relying on NOMA to reduce latency, we aim at minimizing the duration of a single training round of the common FL model, by jointly optimizing the duration of the EH phase, the durations of the EHUs? uplink transmissions of the EHUs, the EHUs’ transmit powers, and the EHUs’ processor frequencies. By transforming the resource allocation problem into a convex optimization problem, we have developed a resource allocation scheme with very low computational complexity.
Federated learning is a new communication and computing concept that allows naturally distributed data sets (e.g., as in data acquisition sensors) to be used to train global models, and therefore successfully addresses privacy, power and bandwidth limitations in wireless networks. In this letter, we study the communications problem of latency minimization of a multi-user wireless network used to train a decentralized machine learning model. To facilitate low latency, the wireless stations (WSs) employ non-orthogonal multiple access (NOMA) for simultaneous transmission of local model parameters to the base station, subject to the users' maximum CPU frequency, maximum transmit power, and maximum available energy. The proposed resource allocation scheme guarantees fair resource sharing among WSs by enforcing only a single WS to spend the maximum allowable energy or transmit at maximum power, whereas the rest of the WSs transmit at lower power and spend less energy. The closed-form analytical solution for the optimal values of resource allocation parameters is used for efficient online implementation of the proposed scheme with low computational complexity.
This paper investigates the occurrence of the double descent phenomenon within the domain of federated learning. It derives a closed-form solution to calculate the mean excess risk in the coefficients of a linear regression model, with the theoretical findings verified through simulations. The results confirm the presence of the double descent effect in the case of federated learning. This effect is especially pronounced in a heterogeneous scenario where local datasets of collaborating users have different sizes and properties. Furthermore, the results unveil unexpected and non-trivial performance behaviors. Interestingly, federated learning even outperforms centralized learning in particular scenarios when the expected excess risk of the former is lower than that of the latter. Finally, our analysis reveals that the typically used random user selection in federated learning exhibits asymptotically lower performance in contrast to the case of excluding only a single carefully chosen user from the learning process.
Background: Bias in medical practice is multifaceted, including treatment variations across race-ethnicity, unconscious bias in healthcare providers' attitudes, and bias in clinical scores. However, far less is known about the potential racial bias in routinely collected, essential information in clinical decision-making, namely vital signs. Research question: Do vital signs embed racial information that can be learned by AI algorithms? Study Design and Methods: Retrospective cohort study of critically ill patients between 2014 and 2015 from the multi-centre eICU-CRD critical care database involving 335 Intensive Care Units (ICU) based in 208 US hospitals, containing 200,859 patient admissions. We extracted 10,763 critical care admissions of patients aged 18 and over, alive during the first 24 hours after admission to ICU with recorded self-reported race as well as at least two measurement of heart rate, oxygen saturation, respiratory rate, and blood pressure. Pairs of racial subgroups were matched based on age, gender, admission diagnosis and APACHE IV scores. Traditional machine learning algorithms, including XGBoost and Logistic regression were used to predict self-reported race using values of vital signs as an input. Results: AI models derived from only six vital signs can predict patients' self-reported race with an AUC of 0.74 (+/- 0.022) between White and Black patients. Technologies used to measure oxygen saturation are a significant source of self-reported racial information (AUC of 0.72 +/- 0.028), in addition to blood pressure measurements (AUC of 0.63 +/- 0.035). Care delivery practices do not present a significant source of racial information (AUC of 0.57 +/- 0.019). However, even when controlling for these known factors, self-reported race can still be learned from vital signs, whose origin we cannot currently explain. Interpretation: Vital signs embed racial information that can be learned by AI algorithms, posing a significant risk to equitable clinical decision-making. Mitigating measures might be challenging, considering fundamental role of vital signs. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This research was partially supported by the WideHealth project - EU Horizon 2020, under grant agreement No 952279. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: All the data for this study is available at https://eicu-crd.mit.edu/gettingstarted/access/ I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All the data for this study is available at https://eicu-crd.mit.edu/gettingstarted/access/
Wearable devices have the ability to generate vast amounts of data that can be put to use in a multitude of applications, particularly in the field of e-health. However, the potential invasion of privacy that comes with utilizing personal data collected by these devices cannot be overlooked. Federated Learning (FL) is a promising solution to this issue that allows models to be trained in a decentralized manner while keeping user data on their own devices. This approach effectively minimizes the risk of privacy breaches and has the potential to be employed in a variety of applications where the protection of user data is of utmost importance. This paper focuses on the use of FL in predicting hospital readmission in 130-US diabetes hospitals for a data set collected over an 9-year period. The results suggest that FL can achieve comparable performance while maintaining privacy and diversity of data. This is an essential aspect of FL, as it enables continuous real-time learning without compromising privacy.
Given the Internet of Things’ rapid expansion and widespread adoption, it is of great concern to establish secure interaction between devices without worsening the quality of their performance. The use of machine learning techniques has been shown to improve detection of anomalous behavior in these types of networks, but their implementation leads to poor performance and compromised privacy. To better address these shortcomings, federated learning (FL) has been introduced. FL enables devices to collaboratively train and evaluate a shared model while keeping personal data on site (e.g., smart homes, intensive care units, hospitals, and so on), thus minimizing the possibility of an attack and fostering real-time distribution of models and learning. This article investigates the performance of FL in comparison to deep learning (DL) with respect to network intrusion detection in ambient assisted living environments. The results demonstrate comparable performances of FL with DL while achieving improved data privacy and security.
The past decade has seen substantial growth in the prevalence and capabilities of wearable devices. For instance, recent human activity recognition (HAR) research has explored using wearable devices in applications such as remote monitoring of patients, detection of gait abnormalities, and cognitive disease identification. However, data collection poses a major challenge in developing HAR systems, especially because of the need to store data at a central location. This raises privacy concerns and makes continuous data collection difficult and expensive due to the high cost of transferring data from a user’s wearable device to a central repository. Considering this, we explore the adoption of federated learning (FL) as a potential solution to address the privacy and cost issues associated with data collection in HAR. More specifically, we investigate the performance and behavioral differences between FL and deep learning (DL) HAR models, under various conditions relevant to real-world deployments. Namely, we explore the differences between the two types of models when (i) using data from different sensor placements, (ii) having access to users with data from heterogeneous sensor placements, (iii) considering bandwidth efficiency, and (iv) dealing with data with incorrect labels. Our results show that FL models suffer from a consistent performance deficit in comparison to their DL counterparts, but achieve these results with much better bandwidth efficiency. Furthermore, we observe that FL models exhibit very similar responses to those of DL models when exposed to data from heterogeneous sensor placements. Finally, we show that the FL models are more robust to data with incorrect labels than their centralized DL counterparts.
AbstractBackgroundracial bias has been shown to be present in clinical data, affecting patients unfairly based on their race, ethnicity and socio-economic status. This problem has the potential to be significantly exacerbated in the light of Artificial Intelligence-aided clinical decision making. We sought to investigate whether bias can be introduced from sources that are considered neutral with respect to ethnicity and race and consequently routinely used in modelling, specifically vital signs.Methodsto perform our analysis, we extracted vital signs from 49,610 admissions from a cohort of adult patients during the first 24 hours after the admission to the Intensive Care Units (ICU), derived from multi-centre eICU-CRD database and single-centre MIMIC-III database, spanning over 208 hospitals and 335 ICUs. Using heart rate, SaO2, respiratory rate, systolic, diastolic, and mean blood pressure, we develop machine learning models based on Logistic Regression and eXtreme Gradient Boosting and investigate their performance in predicting patients’ self-reported race. To balance the dataset between the three ethno-races considered in our study, we use a matching cohort based on age, gender, and admission diagnosis.Findingsstandard machine learning models, derived solely on six vital signs can be used to predict patients’ self-reported race with AUC of 75%. Our findings hold under diverse patient populations, derived from multiple hospitals and intensive care units. We also show that oxygen saturation is a highly predictive variable, even when measured through methods other than pulse oximetry, namely arterial blood gas analysis, suggesting that addressing bias in routinely collected clinical variables will be challenging.Interpretationour finding that machine learning models can predict self-reported race using solely vital signs creates a significant risk in clinical decision making, further exacerbating racial inequalities, with highly challenging mitigation measures.FundingThe funders had no role in the design of this study.
Wireless network (radio) virtualization and its synergy with ML/AI-based technologies is a novel concept that can efficiently address problems of legacy networks, such as flash crowds. This paper discusses the integration aspects of intelligence-based technologies with Sate-of-the-Art end-to-end reconfigurable, flexible and scalable network architecture, capable of handling demands in flash crowd scenarios. The presented results, demonstrate that advanced solutions based on ML can significantly improve the network proactivity and adaptivity by reliably predicting flash crowd scenarios. The results also show that in case of low dataset fidelity, conventional statistical models are a more suitable option.
An important target for machine learning research is obtaining unbiased results, which require addressing bias that might be present in the data as well as the methodology. This is of utmost importance in medical applications of machine learning, where trained models should be unbiased so as to result in systems that are widely applicable, reliable and fair. Since bias can sometimes be introduced through the data itself, in this paper we investigate the presence of ethnoracial bias in patients' clinical data. We focus primarily on vital signs and demographic information and classify patient ethnoraces in subsets of two from the three ethnoracial groups (African Americans, Caucasians, and Hispanics). Our results show that ethnorace can be identified in two out of three patients, setting the initial base for further investigation of the complex issue of ehtnoracial bias.
The original version of this book was inadvertently published with the first author surname incorrect, which has now been corrected.
This paper presents the generic definitions, basic functionalities and current research trends in Cloud Radio Access Networks and its derivatives, Virtual Radio Access Networks and Open Radio Access Networks. Moreover, the paper provides practical results, insights, and lessons learned regarding the limitations and unforeseen issues of Radio Access Networks virtualization. The paper also discusses the potential developments and possibilities for commercial roll out of novel Radio Access Networks approaches.
The FALCON project, presented in this paper, focuses on dynamic on-demand virtual resource allocation for wireless network environments, leveraging a highly flexible and adaptable wireless network to provide reliable and efficient communications in flash crowd scenarios and emergency situations.
End-to-end network slicing is a novel concept based on virtualization and softwarization technology that can efficiently address the problems of legacy networks. It can leverage agile physical and network layer adaptability, and foster optimal user and system performance. These features make end-to-end network slicing suitable for scenarios such as flash crowds and emergency situations. This article presents an end-to-end network slicing framework capable of addressing the demands of flash crowd and emergency scenarios. The article provides details about these scenarios, and it introduces an end-to-end agile, flexible, and scalable wireless system architecture consisting of softwarized components that can be orchestrated to fulfill the underlying system requirements. By addressing important key performance indicators with experimental measurements, the article demonstrates the applicability of our slicing framework for flash crowd scenarios.