Gradient boosting models predict retail demand effectively, yet inventory planners distrust predictions they cannot interpret. This paper addresses the trust gap through a four-level Explainable AI (XAI) framework that pairs LightGBM prediction with layered interpretability: Permutation Feature Importance (PFI) for global ranking, Partial Dependence Plots (PDP) for marginal effects, SHapley Additive exPlanations (SHAP) for instance-level attribution, and Individual Conditional Expectation (ICE) plots for heterogeneity detection. ICE analysis, the key extension beyond prior work, reveals when averaged trends mask divergent product behavior, justifying segment-specific inventory policies. Validated on the UCI Online Retail II dataset with 41 engineered features, the proposed model achieves RMSE of 33.71 and R(2 )of 0.243 on inherently intermittent stock-keeping unit (SKU)-level daily demand, significantly outperforming most baselines (paired t-test, p <0.001; CatBoost not significant, p = 0.917). Ablation experiments confirm rolling statistics as the dominant feature group ( Delta R-2=-0.064 upon removal). ICE heterogeneity analysis identifies product clusters where temporal features produce opposite effects, demonstrating that uniform reorder policies systematically misallocate inventory for specific segments.
Modern software development relies on automated build systems that compile and test code whenever developers make changes. Predicting whether these builds will succeed or fail before execution could save computational resources and developer time. However, many machine learning models for build prediction suffer from temporal data leakage, a methodological flaw where the model inadvertently uses information that would only be available after the build completes, producing artificially inflated accuracy that fails in real-world deployment. This study develops a three-type taxonomy to systematically identify and prevent such leakage: (1) Direct Outcome Encoding (using the build result itself as a feature), (2) Execution-Dependent Metrics (information generated during build execution), and (3) Future Information Leakage (using data from chronologically later builds). Applying this taxonomy reveals that prior studies reporting 95-99% accuracy likely used contaminated features, while realistic accuracy is substantially lower. The methodology is validated on 175,706 builds from two open-source CI/CD platforms spanning 10 years: TravisTorrent (100,000 builds, 2013-2017) and GHALogs (75,706 workflows, 2023). Removing leaky features reduces accuracy by 15.07 percentage points on TravisTorrent (97.8% to 82.73%) but only 0.48 points on GHALogs (83.77% to 83.30%), revealing that modern GitHub Actions' tight integration with repositories enables accurate prediction from static project metadata alone. Using only legitimately available pre-build features, Random Forest classifiers achieve 82.73% (TravisTorrent) and 83.30% (GHALogs) accuracy, sufficient for practical deployment. Surprisingly, project maturity and build history prove more predictive than code complexity metrics, suggesting organizational factors outweigh code quality. The models generalize across programming languages (Java, Ruby, Python, JavaScript) with minimal performance variation. Open-source tools for detecting temporal leakage in any software prediction task are provided.
Retail marketing measurement increasingly requires granular campaign-level insights without relying on user-level tracking. However, the two dominant approaches, Marketing Mix Modeling (MMM) and Multi-Touch Attribution (MTA), often produce fragmented insights. MMM is privacy-safe and robust for channel-level planning but is too coarse for campaign optimization, while MTA provides granular attribution but has become less reliable under increasing privacy restrictions. We propose Integrated Marketing Attribution (IMA), a unified framework that combines MMM with channel specific Bayesian attribution models to derive campaign-level effects from aggregated data. By leveraging MMM-informed priors, IMA delivers granular, privacy-safe attribution while preserving consistency with MMM.
Due to the changing nature of online transactions and the growing sophistication of cybercriminals, the fast expansion of digital banking systems has greatly raised the demand for efficient fraud detection. More sophisticated systems must be developed since conventional fraud detection techniques frequently fail to meet these demands. In order to identify fraud in real-time financial transactions, this research suggests a hybrid deep learning model that blends Convolutional Neural Networks (CNNs) with Long Short-Term Memory (LSTM) networks. LSTMs describe temporal dependencies, allowing the system to comprehend both the individual transaction characteristics and their sequences across time, whereas CNNs are utilized to extract spatial aspects from transaction data. The model is evaluated using publicly accessible financial transaction datasets and contrasted with conventional machine learning algorithms, such as support vector machines and decision trees. According to experimental data, the hybrid CNN-LSTM model performs better than traditional techniques in terms of computing efficiency, recall, and detection accuracy. The results show that the suggested model is a useful tool for digital banking systems since it is especially good at spotting fraudulent activity in real time. This strategy gives financial institutions a strong way to protect transactions in a world that is becoming more and more digital by offering notable advancements in the early detection of fraud.
Artificial intelligence-based autonomous decision systems are being implemented in the fields of finance, healthcare, transportation, and governance of the people. They are used to improve the accuracy, speed, and scalability of decisions through processing high volumes of data and autonomously generating responses without requiring human intervention. Nevertheless, the adoption of autonomous decision systems is rapid, and it presents new governance issues in transparency, accountability, algorithms, as well as operational risks. The old models of governance still depend on manual supervision and fixed policy enforcement that cannot be used to deal with the dynamic AI-driven environments. The following paper proposes a risk-based governing framework that is aimed at providing responsible and trustworthy functioning of autonomous decision systems. The suggested framework incorporates risk identification, risk assessment, governance policy enforcement, and on-going monitoring mechanisms to handle operational and ethical risks. The framework enables organizations to implement the right governance controls by categorizing autonomous systems according to the level of risk and still providing flexibility in operations. Moreover, the architecture has elements of explainability and compliance monitoring to reinforce transparency and accountability. The research emphasizes the role of risk-based governance in allowing organizations to reduce the risks associated with decisions, enhance the adherence to regulations, and increase trust in AI-based decision systems. The suggested framework offers an organized method of governance that facilitates innovation and sustainable use of autonomous decision technologies.