
Transaction fraud detection remains difficult because fraudulent events are highly imbalanced, temporally drifting, cost asymmetric, and often supported by limited auditability in model decisions. Existing tabular, tree-based, and deep fraud detectors can achieve strong predictive performance, but they often provide weak verifier-grounded reasoning, incomplete probability calibration, and limited integration with cost-sensitive decisioning. This paper proposes TEMPLAR-Fraud, a verifier-grounded, calibrated, and cost-sensitive framework for transaction fraud detection. The framework combines SAINT-based transaction representation learning, CatBoost triage, routed symbolic template reasoning, deterministic rationale verification, Bayesian stacking, one-dimensional optimal transport calibration, and validation-selected cost-sensitive thresholding. Evaluation was conducted on BAF Base and IEEE-CIS Fraud Detection using chronological train/validation/test splits, internal testing, same-dataset temporal future-slice evaluation, calibration analysis, cost-sensitive utility assessment, robustness testing, adversarial stress testing, computational-efficiency measurement, and explanation-safety analysis. On the internal chronological test sets, TEMPLAR-Fraud achieved AUROC/AUC-PR/F1 scores of 0.918/0.498/0.557 on BAF Base and 0.972/0.701/0.747 on IEEE-CIS Fraud Detection. Under same-dataset temporal future-slice evaluation, it retained AUROC/AUC-PR/F1 scores of 0.901/0.452/0.518 and 0.951/0.642/0.687, while reducing ECE after calibration to 0.014 and 0.013 and producing the highest reported net savings under the stated experimental cost model. These findings suggest that framework-level integration of routed prediction, verified rationale generation, calibration, and cost-sensitive thresholding can improve fraud-detection evaluation under controlled benchmark protocols.
Cybercrime research in Nigeria has largely focused on offender motivations, victim experiences, and regulatory gaps, with limited use of official administrative data to examine the institutional structure of cybercrime enforcement. This study provides a systematic analysis of cyber-related convictions recorded by the Economic and Financial Crimes Commission (EFCC) in 2022. Using a keyword-based content analysis of publicly available EFCC conviction records (N = 3785), the study identifies and classifies cyber-enabled offences and develops an institutional typology of prosecuted cybercrime in Nigeria. A total of 524 cyber-related cases were extracted from the full conviction dataset and grouped into analytic offence categories. These categories were inductively derived from the offence descriptions recorded in the dataset and include: impersonation and identity fraud; cybercrime and affiliated offences (general); computer-related offences; and business email compromise. Descriptive statistical analysis was used to examine the distribution of offence types and to assess patterns in enforcement focus. Findings indicate that Nigerian cybercrime enforcement is heavily concentrated on impersonation-based and cyber-affiliated offences, with relatively limited representation of technically driven cyber offences. The study demonstrates how institutional data can be used to generate empirical insights into national cybercrime profiles and highlights structural concentration in enforcement priorities. The results contribute to cybercrime scholarship by shifting attention from perception-based studies toward evidence derived from formal criminal justice records. Importantly, the study’s scope is limited to EFCC conviction records, and its findings reflect the enforcement priorities of this specific agency rather than the full landscape of cybercrime prosecution in Nigeria.
Breast cancer remains a leading cause of mortality among women in low- and middle-income countries (LMICs), compounded by fewer radiologists available and resources for diagnosis. A narrative review that summarizes 45 peer-reviewed publications to date from 2018 to 2025 is presented for deep learning (DL) models for mammography detection for breast cancer, targeting low-resource-critical architectures for development in LMIC settings. We compare convolutional neural networks (CNNs), hybrid CNN-support vector machine (SVM) models, recurrent/LSTM networks, and lightweight architectures such as MobileNet and EfficientNet. In quantitative synthesis the reported diagnostic accuracy for MobileNet is between 89 and 92
In many modern applications, analysts observe only aggregated versions of discrete data, such as intervals or fixed-bin frequency summaries built from latent counts. Within the Symbolic Data Analysis framework, these aggregates are symbols generated by a known mapping of unobserved micro-data. Building on the general symbolic-likelihood perspective for aggregated data, this paper develops a likelihood-based treatment tailored to symbolic data generated from latent discrete count models. We construct the symbolic likelihood by summing the latent model over all configurations compatible with each observed symbol, and we derive a score identity showing that the symbolic score equals the conditional expectation of the complete-data score given the symbol. For discrete exponential-family models, this leads to simple estimating equations and EM-type updates that mirror their full-data counterparts; in the Poisson fixed-bin frequency case, the resulting fixed-point iteration is particularly tractable. A simulation study quantifies the efficiency loss induced by different symbol designs, showing that fixed-bin frequency symbolic maximum likelihood estimators recover most of the information in the full data, while min–max intervals and midpoint heuristics can perform noticeably worse. The methodology is illustrated with NBA free-throw attempt data, where only season-level grouped-count summaries of game-level counts are used. Symbolic Poisson and Negative Binomial models are fitted and compared, and goodness-of-fit is assessed entirely in the symbol space via Pearson-type statistics and parametric bootstrap.
European grapevine (Vitis vinifera L.) is a climate–sensitive perennial whose flowering and ripening govern yield and quality. Phenological records from monitoring programs are typically collected at irregular intervals, so true transition dates are interval-censored, and many site-years are right-censored. We develop a reproducible workflow that treats phenology as a time–to–event outcome: Status Intensity observations from the USA–NPN are converted to interval bounds, linked to NASA POWER daily weather, and analyzed with parametric accelerated failure time (AFT) models (Weibull and log–logistic). To avoid outcome–dependent bias from aggregating weather up to the event date, antecedent conditions are summarized in fixed pre–season windows—DOY 1–120 before flowering and DOY 1–180 before ripening—and standardized; quality–control filters require at least 70
Clinical decision-making increasingly relies on data-driven tools, but most systems today are still predictive models that work at isolated time points. Reinforcement learning (RL) provides a different approach by optimizing sequences of actions under uncertainty. It’s often seen as a foundation for more “agentic”AI in healthcare. We conducted a systematic literature review of RL-based clinical decision support systems (CDSS) published between 2020 and January 2026. We reviewed 66 studies, looking at the clinical domain, decision type, RL methods, data, and system maturity. RL-based CDSS are mostly used in critical care, cardiology, oncology, and diabetes, focusing on therapeutic dosing optimization. Actor-critic and policy-gradient methods are mainly used in continuous physiological/device-control settings. Most systems are trained offline using historical data: 66.7
Abstract This review synthesizes research on machine learning algorithms effectiveness in detecting discriminatory patterns in mortgage lending, focusing on decision trees, neural networks, and support vector machines, to address challenges in bias identification and mitigation. The review aimed to evaluate algorithmic effectiveness, benchmark bias detection and mitigation approaches, analyze fairness metrics, compare model interpretability, and examine challenges in addressing intersectional biases. A systematic analysis of empirical studies primarily from the United States employing diverse mortgage datasets and fairness evaluation frameworks was conducted. Findings indicate that decision trees offer high interpretability and effective bias detection, while neural networks and support vector machines achieve superior predictive accuracy but suffer from low transparency. Fairness metrics reveal persistent biases despite high accuracy, and bias mitigation techniques such as reweighting and fairness constraints reduce discrimination but often entail trade-offs with model utility. Intersectional fairness remains underexplored, with limited integration of multi-attribute bias assessments. Overall, balancing accuracy, fairness, and interpretability remains a critical challenge, with explainable AI methods enhancing transparency in complex models. These findings underscore the need for standardized fairness evaluation, improved intersectional bias detection, and transparent algorithmic design to advance equitable mortgage lending practices.
Abstract This study investigates automatic dialect identification for the Somali language, focusing on its two primary dialects: MAXAA TIRI and MAAY. Somali exhibits substantial dialectal variation, which poses challenges for natural language processing (NLP) applications in low-resource settings. To support dialect-aware NLP research, we construct and manually annotate a dataset of 8947 Somali text samples collected from heterogeneous sources, including social media, news outlets, blogs, and formal documents. The study evaluates a range of traditional machine learning and deep learning models, including Naive Bayes, Support Vector Machines (SVM), and Bidirectional Long Short-Term Memory (BiLSTM) networks, for dialect classification. Experimental results show that Naive Bayes and BiLSTM achieve high classification performance under controlled evaluation settings. To mitigate overfitting and source bias, we apply source-aware data splitting, duplicate removal, and ablation analyses. However, results should be interpreted in light of dataset construction constraints, including expert-assisted translation for portions of the MAAY data. This work contributes a linguistically validated Somali dialect dataset and provides empirical insights into the effectiveness of machine learning and deep learning approaches for dialect identification in low-resource contexts.
Abstract Objectives The objective of this dataset is to provide multi-stakeholder survey data collected across enterprises, academia, and government within South Africa’s medical device innovation system. The data was collected to provide data-driven insights into the country’s medical device innovation system, addressing the urgent need for local medical device development due to inappropriate imported devices. This survey is a critical input into a broader goal: the development of a decision support framework to guide local medical device development in South Africa. Data description The dataset comprises multi-stakeholder survey responses, organized in two primary Excel sheets per stakeholder group namely: ‘Variables’ (questionnaire items) and ‘Data’ (raw, anonymized responses). It includes 90 responses (61 completed, 29 partial) from government (53.8% response rate), enterprises (27.2% response rate), and academia (60% response rate). This data is valuable for understanding stakeholder perceptions, demographic information, and insights into key innovation activities within the South African medical device Technology Innovation System, offering a robust foundation for analysing operational dynamics, challenges, and opportunities.
The complexity and difficulty of the ongoing and unstoppable cybercrimes in the traditional or conventional Artificial Intelligence (AI) system create the worst problems for the recent day’s modern cyberspace. The traditional systems are always dependent on centralized systems, which have insufficient performance to prevent and reveal DDoS attacks and modern cybersecurity challenges. The traditional systems of artificial intelligence (AI), machine learning (ML), and Deep Learning (DL) are limited only to the task of detecting DDoS attacks; they are not able to protect the client–server network cyberspace from DDoS attacks. Centralized-based property of the traditional systems produces a particular failure that makes them more attractive victims in cyberspace by DDoS attack criminals. The difficulty of deploying Artificial Intelligence (AI) cybersecurity techniques often results in poor DDoS attack mitigation performance, leaving organizations vulnerable and susceptible to DDoS attacks. The principal and major goal of the study is to formulate a strong Blockchain Technology-based Cyber Security emerging technology-oriented technique, particularly designing and modeling for protecting client–server network cyberspace from DDoS-based cyber-attacks. We have a few numbers of system methodologies like data collection (i.e., related literature review, dataset selection and collection, and experimental data collection), selection of performance evaluation parameters, validation process, experimental equations of performance evaluation metrics, and selection and collection of hardware and software resources. Different models are trained and tested on the CIC-DDoS2019 dataset; we have seen that there are different values of performance evaluation metrics for different models. The CNN model showed the value of 98.5
Recirculating aquaculture systems (RAS) and biofloc technology (BFT) are increasingly recognized as sustainable aquaculture approaches, as they enhance water-use efficiency, improve biosecurity, allow high production intensity, and mitigate environmental impacts by reducing effluent discharge and promoting internal nutrient recycling. The present study aims at thescientometric analysis of emerging intensive aquaculture technologies, especially Recirculating Aquaculture System (RAS) and Biofloc Aquaculture System (BFT), in terms of examination of the publication trend(s), citation pattern, prolific contributors, underpinning the research hot-spots and the top journals publishing articles on the subject under study. The data was retrieved from ‘Scopus’ and downloaded data file was processed in ‘Biblioshiny’ application of ‘R’ software for scrutinizing records and standardizingthe data where required. The study encompasses a total of 853 articles that have been published in 369 journals, with an average annual growth rate of 1.45
Recommender system applications in the financial sector, specifically banking, in fast growing cities like Accra are characterized by serious challenges such as the use of stagnant distance-based search, unreliable GPS positioning, unreliable geocoding accuracy, routes oblivious to congestion, and lack of branch-scaled level of service availability verification. For instance, customers can be referred to the closest bank brach that is located geographically, but is not offering the service that is being requested or even not accessible practically based on the traffic. Such restrictions augment the risk of misguiding and inconveniencing customers. This paper proposes a framework of Geo-Banking, which is a geospatial-based, location-aware recommender system that encompasses spatial indexing, route optimization through networks, service validation and location-awareness systems under a single architecture. Geo-Banking utilizes R-tree indexing to provide efficient retrieval of candidates and congestion-aware routing with the shortest paths to optimize travelling time instead of Euclidean distance. The extra validation layers eliminate false positive geographic identifications and enhance reliability of recommendations. Empirical performance in the Accra Metropolitan Area also indicates better accuracy and recall of service similarity, less geographic mislocation and better computation efficiency than the baseline models. Ablation analysis also supports the fact that GIS-based spatial indexing, as well as routing, adds to the performance improvements over the traditional recommender logic. The Geo-Banking framework is therefore a real-time and scalable urban financial service recommendation model that can be used in smart-city implementation.
We present a dataset designed to evaluate the capacity of Large Language Models (LLMs) to generate context-aware simulations of urban experiences across diverse cultural and socioeconomic settings. The dataset provides a globally distributed and scalable resource for the systematic analysis of daily routines, mobility patterns, and social interactions in major cities. The dataset comprises 1260 simulated urban experience narratives generated by three Large Language Models (LLMs) across 21 major cities in seven global regions. Each narrative captures a structured weekend scenario for one of four predefined actor profiles, representing distinct socioeconomic roles within specific city contexts. A standardized and iteratively refined prompt framework was applied to ensure comparability across cities, actor profiles, and models while preserving controlled generative variability. Actor-level attributes and city-specific contextual information were incorporated to enhance contextual grounding. Additional quality control procedures were implemented to reduce redundancy and maintain temporal and spatial coherence within individual simulations. The resulting corpus constitutes a scalable and systematically structured resource for comparative urban analysis. It enables examination of how contemporary LLMs encode and reproduce representations of urban environments across diverse cultural and socioeconomic settings, supporting the assessment of narrative coherence, cross-model variability, and potential representational bias. As such, it establishes a structured basis for systematic comparison and methodological evaluation of LLM-generated urban narratives.
Software is indispensable for implementing stated preference (SP) studies. Many software packages, particularly free and open-source ones, operate primarily through commands; this can create challenges for students learning SP methods, as those without prior programming experience may face both the psychological burden and the time cost of simultaneously learning SP methods and skills of the associated programming. To support the teaching and learning of introductory-level SP methods, this study developed an R-based graphical user interface (GUI) series designed to conceal coding barriers that may hinder students’ ability to learn the methods. This paper provides lecturers with an overview of SP methods to help them grasp the functionality and features of the SP-GUI series, enabling them to assess its usefulness for teaching SP methods. The SP-GUI series is freely available as R Commander plugin packages and focuses on the main groups of SP methods: dichotomous choice contingent valuation, discrete choice experiments, and three types of best–worst scaling. These packages offer functionalities for designing choice sets, collecting responses, estimating models, and conducting post-estimation analyses. Consequently, students can learn R commands effectively without needing coding skills owing to the dependency of the GUI series on R Commander.
Abstract Brain age estimation using magnetic resonance imaging is a promising biomarker for detecting accelerated aging and neurodegenerative disorders. However, the development of robust clinical models is severely hampered by the “Effects of Site”, where scanner-specific biases obscure biological signals in multi-center datasets. In this study, we propose a novel harmonization strategy, Inter-Site SMOTE, which generates synthetic training data by interpolating between age- and gender-matched participants from different sites. We hypothesize that these synthetic samples populate the sparse regions between site distributions, effectively bridging domain gaps while preserving biological integrity. We systematically evaluated this approach using four large neuroimaging datasets ( $$N=2031$$ ) in a leave-one-site-out regression task. Our results demonstrate that Inter-Site SMOTE significantly improves generalization to unseen scanners compared to standard data pooling. Crucially, we show that standard statistical harmonization (ComBat) fails to improve predictive performance in this setting due to inference-time assumptions, whereas our data-centric approach enhances robustness. Furthermore, we provide empirical evidence that the improvement is driven by the specific geometry of cross-site interpolation, as intra-site augmentation failed to yield comparable gains. This work presents a simple, effective solution for multi-site harmonization that circumvents the limitations of statistical adjustment methods, paving the way for more generalizable prediction models.
Leukemia is recognized as a type of blood cancer originating in the bone marrow and is considered a major global health concern resulting from the uncontrolled proliferation of immature white blood cells in the bloodstream. A mathematical model was developed to evaluate leukemia treatment using CAR T-cell therapy and chemotherapy. The system, expressed through ordinary differential equations, integrates adoptive T-cell infusion and chemotherapeutic effects to assess their influence on leukemia progression. Stability and numerical analyses show that the model becomes stable when immune stimulation exceeds a critical threshold. T-cell infusion significantly reduces cancer and infected cell levels, while higher antigenicity enhances immune efficiency. Chemotherapy dynamics were analyzed under immune and drug interactions, emphasizing the roles of tumor mortality, drug decay, and cytotoxic T-lymphocyte (CTL) activation. Comparative results indicate that CAR T-cell therapy provides a stronger and longer-lasting reduction in leukemia cells than chemotherapy alone. Graphical simulations provide additional insight into the distinct roles of each treatment, supporting the development of optimized therapeutic strategies.
Collecting and analyzing data are critical abilities that help students learn before they get their first job. However, how lecturers or teachers can encourage students with diverse interests to pursue such data science skills more successfully remains unclear. We conducted a case study on data science education at the University of Tsukuba, involving a total of 8,509 participants between 2019 and 2023. Our experiment shows that preparing introductory videos related to students’ respective colleges increased the objective test score by 16.4
A sum-wise formulation is proposed for the Kaplan–Meier product limit estimator of partially right-censored survival data. The population estimator is expressed as a sum over the individual units’ empirical and semi-empirical contributions, for observed and censored failures, respectively. This intuitive decomposition is applied to visualize the different contributions of failed and censored units to the overall population estimator.
Data science has rapidly evolved over the past decade, emerging as a transformative discipline at the intersection of computational methods, machine learning, and big data analytics. The exponential growth of publications highlights the need for a systematic evaluation of global research trends. This study aims to map and analyze a decade of data science research (2015–2025), identifying key contributors, thematic trends, and collaboration patterns, while applying classical bibliometric laws such as Lotka’s and Bradford’s. Scientometric approach was employed using data retrieved from the Scopus database, covering 27,108 documents. Tools such as Biblioshiny and VOSviewer facilitated the analysis of publication trends, author productivity, keyword co-occurrence, and country collaborations. Metrics like Relative Growth Rate (RGR), Doubling Time (DT), and thematic mapping provided in-depth insights. Findings reveal a strong upward trajectory in publications, with the USA, China and India leading global research output. Zhang, Yilong and Wang, Jianyu emerged as highly prolific authors, while “machine learning” and “artificial intelligence” dominated as research themes. Collaboration networks demonstrate increasing internationalization, though disparities remain in global participation. This study uniquely integrates classical bibliometric laws with advanced visualizations, offering a holistic perspective on data science research dynamics and identifying emerging frontiers for future exploration.
As AI systems continue to increase their capabilities of performing human tasks there is a growing need to understand how the AI system determined its decisions. Interpretability is a concept in trustworthy AI research that is focused on understanding of the inner workings and decisions that come from the AI system. Our previous research revealed that AI developers lack consistent approaches or tools for implementing interpretability. There has been substantial theoretical interpretability research, yet the development of practical approaches in the form of tools to assist AI developers on interpretability remains underexplored in the research. This paper develops a taxonomy of AI interpretability techniques based on an analysis of the research literature using the survey of surveys method across 70 papers, examining 30 papers from 2019 to 2023. This paper develops a hierarchical taxonomy at a lower-level of abstraction to present and build upon relevant research literature. This paper also provides an approach as a practical tool in the form of decision trees to help AI developers identify and categorize interpretability techniques based on applicability and characteristics. Unlike existing interpretability taxonomies, this study introduces AI developer-oriented sample decision trees to aid in operationalizing the selection of interpretability techniques. This bridges theoretical research with practical implementation for AI developers. Preliminary testing and validation with AI developers was conducted to ensure applicability of the proposed taxonomy and decision trees. Preliminary validation demonstrated conceptual clarity and practical relevance, forming a foundation for future large-scale evaluation. This research contributes to further bridge the gap between AI systems and ensuring the practical implementation of interpretability.