Recent developments in big data analysis, machine learning, Industry 4.0, and IoT applications have enabled the monitoring and processing of multi-sensor data collected from systems, allowing for the prediction of the "Remaining Useful Life" (RUL) of system components. Particularly in the aviation industry, Prognostic Health Management (PHM) has become one of the most important practices for ensuring reliability and safety. Not only is the accuracy of RUL prediction important, but the implementability of techniques, domain adaptability, and interpretability of system degradation behaviors have also become essential. In this paper, the data collected from the multi-sensor environment of complex systems are processed using a Functional Data Analysis (FDA) approach to predict when the systems will fail and to understand and interpret the systems' life cycles. The approach is applied to the C-MAPSS datasets shared by National Aeronautics and Space Administration, and the behaviors of the sensors in aircraft engine failures are adaptively modeled with Multivariate Functional Principal Component Analysis (MFPCA). While the results indicate that the proposed method predicts the RUL competitively compared to other methods in the literature, it also demonstrates how multivariate Functional Data Analysis is useful for interpretability in prognostic studies within multi-sensor environments.
Graph-based learning methods have become increasingly prominent due to their strong performance across diverse applications. Among these, recent frameworks grounded in diffusion processes provide a unifying perspective that extends traditional graph neural network formulations while addressing limitations of standard message-passing mechanisms. Despite these advances, concerns remain regarding the fairness of such models, as they may propagate or amplify biases present in the data. In this work, we introduce a fairness-aware adaptation of graph-based diffusion by modifying the underlying Laplacian operator. Our approach incorporates multiple complementary transformations, including subspace projections, spectral adjustments, and frequency-based filtering, to mitigate bias-related components. Leveraging the intrinsic smoothing properties of graph diffusion, we provide a principled analysis of the resulting behavior and establish theoretical insights into fairness properties. We evaluate the proposed framework on both synthetic and real-world datasets, demonstrating that it achieves competitive performance while improving fairness metrics with limited additional computational cost.
Detecting outliers in Functional Data Analysis is challenging because curves can stray from the majority in many different ways. The Modified Epigraph Index (MEI) and Modified Hypograph Index (MHI) rank functions by the fraction of the domain on which one curve lies above or below another. While effective for spotting shape anomalies, their construction limits their ability to flag magnitude outliers. This paper introduces two new metrics, the Area-Based Epigraph Index (ABEI) and Area-Based Hypograph Index (ABHI) that quantify the area between curves, enabling simultaneous sensitivity to both magnitude and shape deviations. Building on these indices, we present EHyOut, a robust procedure that recasts functional outlier detection as a multivariate problem: for every curve, and for its first and second derivatives, we compute ABEI and ABHI and then apply multivariate outlier-detection techniques to the resulting feature vectors. Extensive simulations show that EHyOut remains stable across a wide range of contamination settings and often outperforms established benchmark methods. Moreover, applications to Spanish weather data and United Nations world population data further illustrate the practical utility and meaningfulness of this methodology.
In high-dimensional survival analysis, effective variable selection is crucial both for model interpretation and predictive performance. This paper investigates Cox regression with lasso and adaptive lasso penalties in genomic datasets where covariates far outnumber observations. We propose and evaluate four weight calculation strategies for adaptive lasso specifically designed for high-dimensional settings: ridge regression, principal component analysis (PCA), univariate Cox regression and random survival forest (RSF)-based weights. To address the inherent variability in high-dimensional model selection, we develop a robust procedure that evaluates performance across multiple data partitions and selects variables based on a novel importance index. Extensive simulation studies demonstrate that adaptive lasso with ridge and PCA weights significantly outperforms standard lasso in variable selection accuracy while maintaining similar or better predictive performance across various correlation structures, censoring proportions (0% to 80%) and dimensionality settings. These improvements are particularly pronounced in highly censored scenarios, making our approach valuable for real-world genetic studies with limited observed events. We apply our methodology to triple-negative breast cancer data with 234 patients, over 19 500 variables and 82% censoring, identifying key genetic and clinical prognostic factors. Our findings demonstrate that adaptive lasso with appropriate weight calculation provides more stable and interpretable models for high-dimensional survival analysis.
The Network Scale-Up Method (NSUM) is an estimation framework that aims to determine the size of hidden or hard-to-reach populations from questions such as “How many people do you know who belong to the target population?” The information collected from these questions is commonly referred to as aggregated relational data (ARD), and the estimation of hidden population sizes using NSUM in ARD has been widely applied to key problems in sociology and public health. Note that this approach has been widely used to estimate the size of populations subject to legal or social restrictions, such as sex workers and drug users, who are typically excluded from the formal census. Although voting intention is not a social stigma, this information has become a privacy-sensitive issue, particularly in polarized political contexts, and thus poses a challenge to determining the share of party support. In this work, we introduce a methodology for estimating vote-intention shares using NSUM techniques on ARD. The methodology involves the design of the indirect survey to collect ARD regarding the voting intention of the survey participant’s contacts, jointly with other auxiliary questions, the processing of the data with appropriate filters to eliminate outliers, the study of sample stratification strategies, and finally, the support share estimation for each political group by using different NSUM techniques. The methodology is applied to estimate voting outcomes in the 2023 Spanish general elections, using the Madrid, Andalusia, and Valencia regions as experimental scenarios. Overall, the resulting estimates are competitive with those published by leading private and public survey institutes, despite using a significantly smaller number of participants.
Machine Learning algorithms are ubiquitous in key decision-making contexts such as justice, healthcare and finance, which has spawned a great demand for fairness in these procedures. However, the theoretical properties of such models in relation with fairness are still poorly understood, and the intuition behind the relationship between group and individual fairness is still lacking. In this paper, we provide a theoretical framework based on Sheaf Diffusion to leverage tools based on dynamical systems and homology to model fairness. Concretely, the proposed method projects input data into a bias-free space that encodes fairness constrains, resulting in fair solutions. Furthermore, we present a collection of network topologies handling different fairness metrics, leading to a unified method capable of dealing with both individual and group bias. The resulting models have a layer of interpretability in the form of closed-form expressions for their SHAP values, consolidating their place in the responsible Artificial Intelligence landscape. Finally, these intuitions are tested on a simulation study and standard fairness benchmarks, where the proposed methods achieve satisfactory results. More concretely, the paper showcases the performance of the proposed models in terms of accuracy and fairness, studying available trade-offs on the Pareto frontier, checking the effects of changing the different hyper-parameters, and delving into the interpretation of its outputs.
The Network Scale-up Method (NSUM) is a relatively recent statistical approach for estimating the prevalence of unknown populations through indirect surveys utilizing information about the respondents' social circles. The popularity of NSUM has increased in recent years due to its ability to uphold privacy and cost-effectiveness. However, the NSUM is not exempt from biases resulting from participants' behavior. In addition, the simpler and most popular NSUM estimators are based on averages, making them sensitive to deviations in the samples, which may cause significant errors. This work aims to study how robust procedures can overcome misreporting, contamination, and deviation due to conditions such as barrier effects, prevalence, skewness, and tail length. Specifically, the central objective of the article is to analyze the statistical robustness of NSUM methods, studying whether these methods are affected by outliers or unusual data. We employ eight robust proposals for each of the two classical NSUM estimators. We analyze robust estimators through simulation experiments using synthetic random networks such as Erd & odblac;s-R & eacute;nyi, Scale Free, and Stochastic Block Model structures to model different degree distributions and community structures with different prevalence levels in contaminated and uncontaminated scenarios. We compare the results of the simulations with real data on COVID-19 indicators in the United Kingdom and voting intention in the Spanish General Elections of 2023. This article shows that the classical NSUM estimators perform poorly in contaminated scenarios, while most of the robust proposals are not considerably affected. However, the performance of some robust NSUM estimators decreases under barrier effects. In addition, we observe that distortions created by small prevalence play an important role in selecting the most suitable robust NSUM estimator. Particularly, the robustification of the Mean of Ratios (MoR) estimator based on the Myriad operator typically exhibits the best performance (for MoR methods) across the various social network structures for different prevalence levels, reducing the estimation error regarding the non-robust methods by up to three orders of magnitude in contaminated scenarios.
Epidemiologists and social scientists have used the Network Scale-Up Method (NSUM) for over thirty years to estimate the size of a hidden sub-population within a social network. This method involves querying a subset of network nodes about the number of their neighbors belonging to the hidden sub-population. In general, NSUM assumes that the social network topology and the hidden sub-population distribution are well-behaved; hence, the NSUM estimate is close to the actual value. However, bounds on NSUM estimation errors have not been analytically proven. This paper provides analytical bounds on the error incurred by the two most popular NSUM estimators. These bounds assume that the queried nodes accurately provide their degree and the number of neighbors belonging to the hidden sub-population. Our key findings are twofold. First, we show that when an adversary designs the network and places the hidden sub-population, then the estimate can be a factor of Ω(√n) off from the real value (in a network with n nodes). Second, we also prove error bounds when the underlying network is randomly generated, showing that a small constant factor can be achieved with high probability using samples of logarithmic size O(log n). We present improved analytical bounds for Erdős-Rényi and Scale-Free networks. Our theoretical analysis is supported by an extensive set of numerical experiments designed to determine the effect of the sample size on the accuracy of the estimates in both synthetic and real networks.
Thanks to the rapid technological advancement across scientific and engineering domains, the acquisition of extensive and complex data has become increasingly feasible. One notable source of such data is Functional Magnetic Resonance Imaging (fMRI), which allows for the observation of brain activity in a dynamic and spatial context. The analysis of fMRI data, which are defined over the complicated geometry of the brain, poses unique challenges and opportunities. The advent of fMRI has significantly impacted neuroscience, allowing researchers to pinpoint brain regions involved in specific cognitive activities. However, the existing statistical methods for Functional Data Analysis are often confined to Euclidean domains. This paper introduces an innovative approach to functional Partial Least Squares regression designed for scalar-on-function regression problems. The use of a differential regularization term to accommodate the functional nature of data makes the proposed model particularly suitable for handling data defined over complex domains.
Time series modeling for forecasting tasks is a complex problem of great interest to the scientific community. Most forecasting techniques are performed point by point. An alternative is to approximate time series with a higher-level (granular) representation, such that instead of forecasting numerical time series, the objective is to model and forecast at the granularity of the information. In this study, we propose a new approach for granular modeling of time series using Fuzzy Cognitive Maps (FCM). This approach builds on a previous approach that proposes time-segment granulation, which is then fuzzily clustered into clusters. Then, the cluster centroids are used as concepts in the fuzzy cognitive map (FCM) to define a time-series forecasting model. Specifically, we propose three improvements: (i) to the membership function for obtaining the clusters, (ii) to the optimization algorithm for obtaining the FCM, and (iii) to the function it uses to perform the forecast. Tests were carried out with time series of different types, for both granular and numerical forecasts. The experiments carried out demonstrate the robustness of the proposed approach in time series of different characteristics (time series with trends, stationarity, and seasonality). The approach introduced in the paper consistently outperforms the original approach and classical methods such as ARIMA. Overall, our granular approach is very effective at capturing the complexity of non-linear or abruptly changing series, as well as non-stationary and complex patterns.
With the rapid growth of data generation, advancements in functional data analysis have become essential, especially for approaches that handle multiple variables at the same time. This paper introduces a novel formulation of the epigraph and hypograph indices, along with their generalized expressions, specifically designed for multivariate functional data (MFD). These new definitions account for interrelationships between variables, enabling effective clustering of MFD based on the original data curves and their first two derivatives. The methodology developed here has been tested on simulated datasets, demonstrating strong performance compared to state-of-the-art methods. Its practical utility is further illustrated with two environmental datasets: the Canadian weather dataset and a 2023 air quality study in Madrid. These applications highlight the potential of the method as a great tool for analyzing complex environmental data, offering valuable insights for researchers and policymakers in climate and environmental research.
Recruiting passive candidates, i.e. individuals not actively seeking jobs but open to compelling opportunities, remains one of the hardest challenges in digital recruitment. Motivated by a real collaboration with an industry partner, we introduce the Independent Halting Cascade (IHC) model: a simple but rich agent-based framework that couples network diffusion with the possibility of halting through job applications. Agents can either recommend vacancies to peers or apply themselves, and incentives increase the likelihood of recommendation, mobilizing otherwise passive candidates. The IHC bridges research on social network diffusion, coordinated task completion, and labor economics by modeling heterogeneous skills, job specificities, and network structures, including homophily. We derive analytical boundaries that characterize diffusion and failure regimes, and we show through simulations that the IHC reproduces the empirical chain-length distributions of Travers and Milgram and Dodds with only coarse calibration. Across synthetic (ER, BA, homophilic) and real networks (SMS, e-mail, Twitter), the IHC achieves comparable or higher success rates than direct-recommendation baselines, while requiring fewer applicants. Our findings suggest that the IHC captures core mechanisms of coordinated task completion, offering both a theoretical contribution and a practical foundation for recruitment systems designed to reach and engage passive candidates.
Interpretability of neural networks (NNs) and their underlying theoretical behavior remain an open field of study even after the great success of their practical applications, particularly with the emergence of deep learning. In this work, NN2Poly is proposed: a theoretical approach to obtain an explicit polynomial model that provides an accurate representation of an already trained fully connected feed-forward artificial NN [a multilayer perceptron (MLP)]. This approach extends a previous idea proposed in the literature, which was limited to single hidden layer networks, to work with arbitrarily deep MLPs in both regression and classification tasks. NN2Poly uses a Taylor expansion on the activation function, at each layer, and then applies several combinatorial properties to calculate the coefficients of the desired polynomials. Discussion is presented on the main computational challenges of this method, and the way to overcome them by imposing certain constraints during the training phase. Finally, simulation experiments as well as applications to real tabular datasets are presented to demonstrate the effectiveness of the proposed method.
In this paper we analyze the well-known 'Anonymous bank' call center dataset from a queueing science viewpoint. For this purpose, fitted distributions for both the inter-arrival and service times as well as for customers patiences are integrated in a simulator to infer quantities of interest related to call centers managerial decisions as waiting times, abandonment rates and queue lengths. In particular, it is shown how a type of Markov renewal process, the Markovian arrival process (MAP), is able to capture some of the characterizing properties of arrivals in a modern call center as overdispersion and positive correlation between arrival counts. The work provides a new inference approach for the MAP based on the count process descriptors and presents new properties concerning the dependence structure of the cumulated number of arrivals in a MAP.
Indirect surveys, in which respondents provide information about other people they know, have been proposed for estimating (nowcasting) the size of a hidden population where privacy is important or the hidden population is hard to reach. Examples include estimating casualties in an earthquake, conditions among female sex workers, and the prevalence of drug use and infectious diseases. The Network Scale-up Method (NSUM) is the classical approach to developing estimates from indirect surveys, but it was designed for one-shot surveys. Further, it requires certain assumptions and asking for or estimating the number of individuals in each respondent's network. In recent years, surveys have been increasingly deployed online and can collect data continuously (e.g., COVID-19 surveys on Facebook during much of the pandemic). Conventional NSUM can be applied to these scenarios by analyzing the data independently at each point in time, but this misses the opportunity of leveraging the temporal dimension. We propose to use the responses from indirect surveys collected over time and develop analytical tools (i) to prove that indirect surveys can provide better estimates for the trends of the hidden population over time, as compared to direct surveys and (ii) to identify appropriate temporal aggregations to improve the estimates. We demonstrate through extensive simulations that our approach outperforms traditional NSUM and direct surveying methods. We also empirically demonstrate the superiority of our approach on a real indirect survey dataset of COVID-19 cases.
Introducción: Seguir la evolución del COVID-19 ha sido esencial para la toma de decisiones sanitarias. Pero ello requiere de cifras fiables de los infectados, muertos y hospitalizados, muchas veces complicadas de tener por múltiples razones. El uso de técnicas de estimación rápidas y fiables, como las encuestas indirectas, es una opción. Objetivos: Presentar el proyecto CoronaSurveys, el cual monitorea el COVID-19 combinando encuestas indirectas con el método de ampliación de red (Network Scale-up Method, NSUM). Metodología: El sistema usa encuestas anónimas indirectas en línea para consultar sobre el COVID-19, y el método NSUM para estimar los casos. Las encuestas están en múltiples idiomas con preguntas para rastrear los casos activos, nuevos y de muertes, entre otros. Resultados: CoronaSurveys está operando desde marzo de 2020, y sigue recopilando datos, con más de cien mil respuestas en la actualidad (millones de muestras). El sistema ha hecho buenas estimaciones para España y el Reino Unido, entre otros países. Conclusión: Este sistema es adecuado en países con infraestructura sanitaria limitada o en situaciones de desconocimiento de la pandemia, como al inicio del COVID-19, porque su coste de implementación es pequeño, requiere dispositivos simples para usarlo, y de pocos participantes para obtener buenas estimaciones.
This study aims to evaluate the performance of Cox regression with lasso penalty and adaptive lasso penalty in high-dimensional settings. Variable selection methods are necessary in this context to reduce dimensionality and make the problem feasible. Several weight calculation procedures for adaptive lasso are proposed to determine if they offer an improvement over lasso, as adaptive lasso addresses its inherent bias. These proposed weights are based on principal component analysis, ridge regression, univariate Cox regressions and random survival forest (RSF). The proposals are evaluated in simulated datasets. A real application of these methodologies in the context of genomic data is also carried out. The study consists of determining the variables, clinical and genetic, that influence the survival of patients with triple-negative breast cancer (TNBC), which is a type breast cancer with low survival rates due to its aggressive nature.
Introduction: Monitoring the evolution of COVID-19 has been essential for health decision-making. But this requires reliable numbers of the infected, dead and hospitalized, often difficult to have for multiple reasons. The use of fast and reliable estimation techniques, such as indirect surveys, is an option. Objectives: To present the CoronaSurveys project, which monitors COVID-19 by combining indirect surveys with the Network Scaleup Method (NSUM). Methodology: The system uses indirect online anonymous surveys to inquire about COVID-19, and the NSUM method to estimate cases. The system offers multi-language surveys with questions to track active cases, new cases, and fatalities, among others. Results: CoronaSurveys has been operating since March 2020, and continues to collect data, with over one hundred responses to date (over one million samples). The system has made good estimates for Spain and the United Kingdom, among other countries. Conclusion: This system is suitable in countries with limited health infrastructure or in situations of lack of knowledge about the pandemic, such as at the beginning of COVID-19, because its implementation cost is small and it requires simple devices to use it and few participants to obtain good estimates.
The nn2poly package provides the implementation in R of the NN2Poly method to explain and interpret feed-forward neural networks by means of polynomial representations that predict in an equivalent manner as the original network.Through the obtained polynomial coefficients, the effect and importance of each variable and their interactions on the output can be represented. This capabiltiy of capturing interactions is a key aspect usually missing from most Explainable Artificial Intelligence (XAI) methods, specially if they rely on expensive computations that can be amplified when used on large neural networks. The package provides integration with the main deep learning framework packages in R (tensorflow and torch), allowing an user-friendly application of the NN2Poly algorithm. Furthermore, nn2poly provides implementation of the required weight constraints to be used during the network training in those same frameworks. Other neural networks packages can also be used by including their weights in list format. Polynomials obtained with nn2poly can also be used to predict with new data or be visualized through its own plot method. Simulations are provided exemplifying the usage of the package alongside with a comparison with other approaches available in R to interpret neural networks.
Carlos Baquero合作论文数Departamento de Informática, Universidade do Minho8