
Remote sensing activities entail modelling long distance dependencies beyond the neighbourhood. The classic CNNs, which are founded on fixed-grid Euclidean representation, suffer the inherent shortcoming of being unable to capture non-local dependency, and an uneven structure of geographical things. Despite the fact that Object-Based Image Analysis (OBIA) have proposed the context-dependent spatial primitives, they are nevertheless constrained by the fixed segmentation boundaries and the use of heuristic graph creation that makes them not adaptable. Graph Neural Networks (GNNs) provide a framework of non-Euclidean relational reasoning, however the current approaches presuppose the fixed graphs and fixed semantics of nodes, and restrict dynamic representation of spatial-temporal changes in remote sensing data. It is proposed in this paper that an end-to-end shift to common frameworks between adaptive spatial primitives and graph topologies should be undertaken. These structures recursively co-construct spatial entities and their association with task specific loss cues to allocate new differentiable layouts of groupings and the consequent emergence of edges modulated by the expanding node representation. This multi-modal and multiscale representation is unified in a way that directly facilitates easy inference, and it can also adapt to temporal change. The shortcomings of the lack of labels and computational access to scaling to feedback loops in both feature extraction and relational reasoning are addressed by these systems and finally lead to the enhancement of semantic coherence and representable form. We present the long-desired paradigm that combines UQ, encourages lifelong learning, and overcome the drawbacks of the current fixed-grid or heuristic-graph-based cycles of RS analytics to move to adaptable, and comparatively minimized spatial and relational representations.
The purpose of this study is to investigate and compare several nonparametric regression approaches, including penalized spline methods, B-splines, and smoothing splines. Applying these techniques to simulated and real datasets, such as Iraqi oil export data, focuses on parameter estimation and identifying optimal knot points for predicting periodic and nonlinear trends. The knot points are selected using generalized cross-validation (GCV) to ensure an accurate fit to the data. For time-series data with nonlinearities and periodic patterns in the response variable, this research employs nonparametric regression with sequential explanatory variables. We research simulated data that exhibit periodic patterns similar to economic cycles, as well as nonlinear data that employs complex equations to model interactions among variables. Simulations were conducted across a range of standard deviations and sample sizes. The efficiency of parameter estimation in these synthetic datasets was quantified using the mean absolute average error (MAME). For the empirical application, the parameters of the nonparametric regression models were estimated using monthly Iraqi oil export data, with the MAME employed as the evaluation metric. The effectiveness of these techniques is further evaluated in forecasting future values by calculating the mean absolute percentage error (MAPE). Among the approaches, the penalized spline consistently achieves the lowest average mean squared error across all standard deviation levels and sample sizes in the simulated data, while also demonstrating robust forecasting performance. In contrast, the smoothing spline outperforms the other methods in terms of parameter estimation accuracy.
The study investigated the adoption of AI in the banking sector using a PCA-based k-means clustering method, drawing on data from the Evident AI Index Rankings. The objective was to identify distinct patterns in banks' integration and use of AI technologies, with an emphasis on talent, innovation, leadership, and transparency. Utilizing PCA for dimensionality reduction, the study distilled the intricate aspects of AI adoption into fundamental components, thereby improving the comprehension of clustering patterns among banks. The k-means clustering identified unique segments within the sector, such as early AI adopters, innovation leaders, and conservative implementers, each exhibiting distinct levels of AI maturity and application focus. These findings provided valuable insights into the competitive landscape of AI utilization in banking, highlighting leading institutions in AI-driven transformation and those encountering adoption challenges. The insights from this analysis offered practical implications for stakeholders, guiding strategies for improved AI integration and competitive positioning. The study emphasizes the significance of data-driven benchmarking tools, such as the Evident AI Index, in assessing and guiding technological evolution across the sector.
Lassa fever, a viral hemorrhagic fever endemic to West Africa, poses significant public health challenges with annual case estimates ranging from 100,000 to 300,000 infections and mortality rates reaching 15-20% in hospitalized patients. Current surveillance systems rely predominantly on passive case detection and laboratory confirmation, often resulting in delayed outbreak identification and response. The complex interplay of environmental, climatic, and demographic factors influencing Lassa fever transmission patterns necessitates sophisticated predictive modeling approaches that can process multiple data streams and identify early warning signals for potential outbreaks.This study aims to develop and evaluate an AI-driven prediction model for Lassa fever outbreaks by integrating evolutionary algorithms and Random Forest for optimal feature selection with ensemble learning to enhance early detection capabilities and support proactive public health interventions. We implemented a hybrid machine learning approach combining genetic algorithms Random Forest for feature optimization with XGBoost for model training. Evolutionary algorithms and Random Forest were employed to identify the most predictive feature subsets, followed by XGBoost model training and validation using stratified cross-validation and temporal holdout testing. The evolutionary algorithm + correlation filter approach achieved exceptional performance with 80.04% accuracy, 61.02% macro precision, and 78.29% weighted F1-score, demonstrating significant improvement over traditional Random Forest feature selection (76.73% accuracy).. The model's high accuracy and interpretability make it suitable for integration into existing public health infrastructure, potentially reducing outbreak response time and improving resource allocation for preventive interventions in endemic regions.
As competition for jobs intensifies, job seekers must focus on crafting descriptions that align well with organizations, making it imperative to have a proper resume. The investigation involves the development of an ML-driven Advanced Application Tracking System (ATS) that reviews resumes and provides in-depth feedback on their quality. It will analyze resumes for factors such as relevance to the job descriptions, keyword optimization, structure, and overall presentation. Some aspects include Resume evaluation, rating resumes as poor, good, or excellent based on preset parameters, and providing specific suggestions for improvement. It will then optimize keywords by identifying essential terms from a job description and validating their inclusion and contextual relevance in a candidate’s resume, ensuring alignment with automated HR screening protocols. By integrating a personality-prediction module, the system would analyze resumes using Natural Language Processing (NLP) and sentiment analysis to predict personality traits that may help employers assess whether the applicant will fit their particular company culture and team dynamics. With appropriate guidance, candidates can tailor their job applications to secure the most suitable employment opportunities. Moreover, this tool, built as a result of machine learning combined with natural processing, simplifies the hiring procedure, leading eventually to better resume quality for hiring candidates and consequently better candidate suitability on the recruiting party's side - all these result in the holistic efficiency of the hiring process.
This study employs Natural Language Processing (NLP) techniques to analyze sentiments expressed in Instagram comments about Prabowo Subianto's inauguration as Indonesia's president. The dataset comprises a rich collection of user-generated comments, meticulously preprocessed with the Sastrawi stemmer tailored for Indonesian. This preprocessing stage includes rigorous text cleaning, stemming, and stopword removal, ensuring that the analysis is based on the most relevant linguistic elements. To accurately classify the sentiment of these comments as positive or negative, a logistic regression model has been trained. The model leverages TF-IDF (Term Frequency-Inverse Document Frequency) for effective feature extraction, enhancing the precision of the analysis. With promising results, particularly in identifying uplifting remarks that celebrate the new president's ascendance, this study underscores the essential role of natural language processing in unraveling public sentiment surrounding pivotal political events. The findings of this research not only shed light on the intricate tapestry of public opinion but also pave the way for future sentiment analysis endeavors within the vibrant landscape of Indonesian social media. The model demonstrates robust accuracy, illustrating its effectiveness in interpreting the nuanced sentiments of digital discourse surrounding significant political milestones.
This article presents OptionMC, a Python package designed for educational purposes that implements Monte Carlo methods for European option pricing. We describe the package's architecture and demonstrate its application through systematic testing against established Black-Scholes analytical solutions. The implementation supports both standard Monte Carlo estimation and variance reduction via antithetic variates, allowing examination of convergence patterns and computational efficiency. Our results suggest that Monte Carlo estimates converge toward analytical solutions as the number of iterations increases, with convergence behavior generally consistent with theoretical expectations. Analysis of parameter sensitivity indicates the package appropriately captures fundamental pricing relationships, including volatility effects, time decay, and moneyness considerations. The distributional characteristics of simulated stock prices and option payoffs align reasonably well with theoretical predictions. While OptionMC primarily serves pedagogical objectives rather than high-performance applications, it offers a transparent framework that may benefit students and researchers seeking to understand the practical implementation of option pricing algorithms through Monte Carlo techniques.
Chronic diseases such as diabetes, stroke, and heart disease are major challenges in the global health system. Data-driven risk prediction for this disease is important for supporting more precise and effective medical decisions. This study aims to evaluate the main factors contributing to the incidence of diabetes, stroke, and heart disease using logistic regression analysis. The data used are from health sources and includes demographic variables, lifestyle factors, and health indicators. Logistic regression was used to identify variables significantly associated with each health condition studied. The model was evaluated using p-value, regression coefficient, and confidence interval to assess the significance of risk factors. The results of the analysis showed that age, high blood pressure, cholesterol levels, and body mass index (BMI) contributed significantly to the risk of diabetes, stroke, and heart disease. Physical activity and alcohol consumption negatively affected the risk, while smoking factors did not show strong significance in the model. These findings confirm that certain lifestyle factors and health conditions significantly affect the risk of chronic disease. The implications of this research can inform data-driven prevention and early intervention strategies in the health sector.
This study presents a bibliometric analysis of research trends in linear regression using data retrieved from the Scopus database. The search query included the keywords linear and regression in the title, abstract, or keywords, and filtered results to final-stage journal articles written in English, categorized under the exact keyword Article, affiliated with institutions in the United Kingdom, and published between 2022 and 2025. The analysis, conducted using VOS viewer and Scopus visualization tools, reveals fluctuations in publication counts, with a notable decline in 2025. Leading journals such as PLOS ONE, BMJ Open, and International Journal of Environmental Research and Public Health contribute significantly to the field. At the same time, University College London (UCL) and the University of Oxford emerge as the most influential institutions. Co-authorship and co-citation analyses indicate strong collaborative networks, particularly in medical and epidemiological research. Additionally, funding bodies like the National Institute for Health and Care Research (NIHR) and UK Research and Innovation play crucial roles in supporting these studies. Despite the widespread applications of linear regression across various disciplines, gaps remain in its methodological advancements, particularly in handling high-dimensional data, non- linearity, and real-time decision-making applications. This study highlights the need for future research to explore more adaptive and robust regression models that integrate machine learning techniques and dynamic real-world scenarios.
Accurate traffic prediction in large cities such as Los Angeles is increasingly necessary as cities expand and more vehicles are added to the roads. Using the METR-LA dataset. Using the METR-LA dataset, this study proposes a hybrid deep learning architecture that combines time and space modeling techniques to improve the accuracy and scalability of traffic flow predictions. The dataset consists of multivariate time series data from 207 loop detectors that record traffic speeds every five minutes with very high resolution. This study evaluates five potential model configurations: Long Short-Term Memory (LSTM), Transformer-based TSFormer, a combination of LSTM and TSFormer, Spatio-Temporal Graph Convolutional Network (STGCN), and a model combining STGCN and TSFormer. The evaluation conducted using three performance metrics Mean Absolute Error (MAE), Mean Squared Error (MSE), and Root Mean Squared Error (RMSE) were used to assess how well each model captures complex temporal and spatial relationships. Our results show that the LSTM+TSFormer hybrid model consistently outperforms all other models across all criteria. This model has the lowest MAE (0.0624) and RMSE (0.1204), meaning it is better at learning patterns that occur over time and patterns that occur rapidly. STGCN-based models are quite good at capturing spatial dependencies, but their performance improves when combined with attention-based TSFormer modules. The hybrid models introduced in this study overcome major limitations, including the narrow receptive range of recurrent networks and the inflexible spatial structures assumed in graph-based methods. This work offers important perspectives for developing forecasting models that are not only accurate and scalable but also transparent and adaptable. Future work may explore dynamic graph construction and multimodal input integration to further enhance adaptability in real-world applications.
Sarcasm is common on social media, yet difficult for machines to interpret. Its meaning often relies on conversational tone, speaker intent or situational contrast—signals not directly visible in plain text. This study investigates how far one can go in sarcasm detection using only classical machine learning techniques and hand-crafted feature engineering, without relying on neural architecture or contextual information. Using a 100,000-comment stratified subsample of the Self-Annotated Reddit Corpus (SARC 2.0), I combine word-level and character-level TF–IDF representations with simple stylistic features such as length, punctuation use, and uppercase ratios. Four classical classifiers are evaluated: logistic regression, linear support vector machines, multinomial Naive Bayes, and random forests. Despite the context-free design, logistic regression and Naive Bayes reach F1-scores of approximately 0.57 on sarcastic comments, demonstrating that classical approaches capture part of the underlying signal. The full code is included for reproducibility.