
Formal predictive analysis remains limited in tobacco-dependent economies, where forecasting has largely relied on ARIMA-type time-series models. While widely used, these models impose linearity assumptions that restrict their ability to capture key structural drivers of production. This limitation is evident in recent studies where an ARIMA (1,1,0) model projected Zimbabwean tobacco yield at 1,511.78 kg/ha for 2023, underestimating the observed yield of 2,278 kg/ha by approximately 50.7% [1]. Export forecasting is even less developed, with most existing studies remaining descriptive rather than predictive. The paper reviewed the literature related to tobacco yield forecasting, agricultural export modelling and the application of Machine Learning in crop prediction and trade intelligence systems. Data from across the fields confirm that the Machine Learning techniques Ridge Regression, Random Forest and Gradient Boosting offer superior results to statistical models. The analysis points out three main gaps. First, Machine Learning methods have not been widely applied to tobacco production in sub-Saharan Africa. Second, there is no formal export forecasting model for Zimbabwe that accounts for its multi-year shipment patterns. Third, there is no integrated framework that jointly modelsyield and exports within a unified decision-support system. These gaps highlight the need for more comprehensive, data-driven approaches to forecasting in the tobacco sector.
The World Wide Web has become one of the most valuable resources for data recovery and information releases since it has the largest collection of data and numerous pages or reports. Advances in web mining are the key to unlocking information on the Internet. There are three types of web mining: web content, web structure, and web usage. One of these classes, Web structure mining is the focus of this research. A significant role in the web mining process is played by web structure mining. This paper discusses the experimental results for Link Based Ranking Algorithms and clarifies Web Mining techniques and certain well-known tactics used in Web structure mining.
In this study, we analyzed average sleep durations across 61 countries to investigate the impact of Daylight Saving Time (DST) practices. We identified key metrics influencing sleep and employed statistical correlation analysis to explore relationships among these factors. Countries were categorized based on DST observance, and visualizations were generated to compare sleep patterns between DST and non-DST regions. Our findings indicate that, on average, countries that observe DST tend to have better sleep durations compared to those that do not. However, a more nuanced pattern emerged when accounting for latitude: DST- observing countries at lower latitudes reported shorter sleep durations than their non-DST counterparts, whereas at higher latitudes, DST-observing countries demonstrated longer aver- age sleep durations. These results suggest that the effectiveness of DST in improving sleep may be moderated by a country’s geographical location.
The integration of Sentiment Analysis (SA) and Natural Language Processing (NLP) assists companies in improving customer service by examining actionable insights from unstructured data sources like social media tweets and customer reviews. RoBERTa acts as a sophisticated transformer-based model that is capable of analyzing complex customer sentiment data. This paper assesses the performance of RoBERTa for sentiment classification while investigating solutions to class imbalance issues, sarcasm detection difficulties, and ethical issues. The discussion proposes optimization techniques for RoBERTa on different data sets to register high accuracy and outline future research directions to improve fairness and explainability. The paper concludes its material by introducing mathematical frameworks as performance metric tools in addition to optimization and evaluation procedures.
This study evaluates the effectiveness of a median sector rotation strategy within the Nikkei 500 component sectors, building on prior research that demonstrated superior risk-adjusted returns by selecting midperforming assets. Unlike traditional momentum-based investing, which focuses on winners or losers, the median strategy systematically reallocates capital to sectors with moderate past performance, reducing volatility while maintaining steady growth. Our findings reveal that quarterly and semi-annual rebalancing optimize returns in Japan, differing from U.S.-based studies where monthly rebalancing was more effective. Unlike buy-and-hold investing, the median strategy tends to outperforms total return and drawdown reduction, making it a viable alternative for public investors. By applying structured sector rotation rather than passive indexing, investors gain exposure to Japan’s strongest industries while mitigating downside risk. The results highlight the strategy’s adaptability across markets and suggest broader applications in global equities, fixed income, and multi-asset portfolios for enhanced portfolio resilience.
Data mining techniques are essential for uncovering patterns and trends in various domains. In this analysis, we examine the aspect ratios of world currencies over time, focusing on historical trends, clustering patterns, and mathematical implications. We identify evolving design standards and their significance by collecting and analyzing web-scraped data from multiple sources. Our discussion is structured into four main areas: (i) currency attributes and dataset structure, (ii) statistical trends in aspect ratios, (iii) clustering of currency designs, and (iv) applications of decomposition techniques. We first define aspect ratios and their functional roles, then analyze their evolution using polynomial regression, hierarchical clustering, and seasonal-trend decomposition. Finally, we summarize the impact of our findings on currency design and standardization efforts.
This paper explored the intellectual and cultural transformations reflected in the Acta Eruditorum, a prominent early modern European scholarly journal published from 1682 to 1735. By analyzing Acta’s temporal dataset encompassing more than 7,000 papers, our study examines the time distribution of contributors’ expertise across six domains: Law, Literature, Science, Mathematics, Politics, and Religion. Key findings include the increasing rise of publications in Science and Mathematics, aligning with Enlightenment ideals and the Scientific Revolution. Furthermore, the study shows a significant decrease in religious contributions, reflecting a broader shift from religious perspectives to a greater emphasis on scientific thought and beliefs. Our data visualization techniques and statistical analysis reveal intriguing parallels and contrasts between the Acta Eruditorum and the French Academy of Sciences, highlight- ing their distinct and complementary contributions to the advancement of knowledge. These findings provide valuable insights into how each institution shaped European intellectual history and fostered the exchange of ideas that propelled the Enlightenment era.
Leonhard Euler, one of the most influential mathematicians of the 18th century, has been accredited for introducing a significant portion of modern mathematics. Since its founding, Euler has been an active member of the Imperial Russian Academy of Sciences in Saint Petersburg. With an emphasis on the era of Leonhard Euler’s influence, this study explores the intellectual output of the Russian Academy of Sciences throughout the 18th century. Scholarly work published by the Academy was collected and analyzed by examining various aspects, including authorship and disciplinary trends. Our primary source is the detailed catalog of the Academy published in the nineteenth century by Paul Heinrich Fuss, the secretary of the Imperial Academy. Mining data in this catalog, our findings reveal key contributors, publication patterns, and the evolution of scholarly focus within the academy. Euler emerges as a dominant figure whose prolific output significantly shaped the academy’s intellectual landscape. The study provides insights into the academy’s development and the additional context of Euler’s groundbreaking work.
This study examines the application of data mining techniques to analyze ancient Roman coin datasets and investigate the extent to which Benford’s Law is exhibited in the numerical values of the coins. Ben- ford’s Law predicts the frequency distribution of the first digits in naturally occurring datasets, and its applicability has been demonstrated across diverse fields. This research aims to explore whether ancient Roman coin values conform to this mathematical phenomenon, providing insights into the au- thenticity and naturalness of the data. By employing data mining methods, we analyze the leading digit distribution in a comprehensive dataset of ancient Roman coins. Additionally, we investigate trends and features within the dataset, such as coin composition, weight and diameter. The findings of this study contribute to the broader understanding of historical numismatic data and the relevance of Benford’s Law in historical datasets.
This paper introduces a data-driven investment strategy that leverages data mining techniques to streamline portfolio allocation. Using historical performance data from 15 sectoral indices of the National Stock Exchange of India, the study applies periodic ranking and clustering methods to identify highperforming indices. At predefined intervals—annually, semiannually, quarterly, and monthly—indexes are re-ranked based on historical returns and reassigned to new groups using pattern recognition techniques. An initial investment of $100 is distributed in three dynamically formed groups, each comprising five indices. Returns within each group are reinvested in the same group for subsequent periods, ensuring systematic portfolio evolution. By integrating data mining principles such as ranking algorithms and periodic reassignment, this strategy offers an intuitive yet computationally efficient approach to portfolio management, demonstrating how financial decision making can be optimized through data-driven insights.
Action Rules are rule based systems that extract actionable patterns which are hidden in big volumes of data. Huge amount of data gets generated from Education sector, Business field, Medical domain and Social Media, in a single day. In the technological world of big data, massive amounts of data are collected by organizations, including in major domains like financial, medical, social media and Internet of Things(IoT). Mining this data can provide a lot of meaningful insights on how to improve user experience in multiple domain. Users need recommendations on actions they can undertake to increase their profit or accomplish their goals, this recommendations are provided by Actionable patterns. For example: How to improve student learning; how to increase business profitability; how to improve user experience in social media; and how to heal patients and assist hospital administrators. Action Rules provide actionable suggestions on how to change the state of an object from an existing state to a desired state for the benefit of the user. The traditional Action Rules extraction models, which analyze the data in a non distributed fashion, does not perform well when dealing larger datasets. In this work we are concentrating on the vertical data splitting strategy using information granules and creating the data partitioning more logically instead of splitting the data randomly and also generating meta actions after the vertical split. Information granules form basic entities in the world of Granular Computing(GrC), which represents meaningful smaller units derived from a larger complex information system. We introduced Modified Hybrid Action rule method with Partition Threshold Rho. Modified Hybrid Action rule mining approach combines both these frameworks and generates complete set of Action Rules, which further improves the computational performance with large datasets.
Artificial intelligence (AI) is transforming the retail industry’s approach to data management and decisionmaking. This journal explores how AI-powered techniques enhance data governance in retail, ensuring data quality, security, and compliance in an era of big data and real-time analytics. We review the current landscape of AI adoption in retail, underscoring the need for robust data governance frameworks to handle the influx of data and support AI initiatives. Drawing on literature and industry examples, we examine established data governance frameworks and how AI technologies (such as machine learning and automation) are augmenting traditional data management practices. Key applications are identified, including AI-driven data quality improvement, automated metadata management, and intelligent data lineage tracking, illustrating how these innovations streamline operations and maintain data integrity. Ethical considerationsincluding customer privacy, bias mitigation, transparency, and regulatory compliance are discussed to address the challenges of deploying AI in data governance responsibly.
The landscape of software development has seen a massive shift in the last few years, with rising use of data-driven methods for making product decisions. One area that has made a significant difference is the integration of machine learning and artificial intelligence technologies to inform software engineering practice, including prioritization of product features. Software product feature prioritization is an essential process directly influencing the competitiveness and success of a product. Traditional techniques, though fundamental, tend to fall short in resolving the intricacies of contemporary software ecosystems. This study delves into the revolutionary potential of machine learning (ML) and artificial intelligence (AI) for improving feature prioritization. An extensive literature survey identifies existing trends and their drawbacks, such as inadequate integrated frameworks and scalability and interpretability issues. The suggested framework integrates heterogeneous sources of data, predictive analytics, natural language processing (NLP), and optimization algorithms to support real-time data-driven decision-making
Customer targeting has become a critical component of modern marketing strategies, driven by advancements in Artificial Intelligence (AI). This paper presents a novel AI-powered customer segmentation framework that integrates K-Means clustering, Principal Component Analysis (PCA), and Random Forest classification to enhance predictive analytics for strategic marketing impact. The rationale for selecting these methods is thoroughly discussed, highlighting their strengths over alternatives like DBSCAN, LDA, and SVM. Additionally, baseline comparisons and experimental evaluations demonstrate the effectiveness of the proposed approach. Real-world e-commerce datasets are leveraged to illustrate the model’s ability to generate granular customer insights. Unlike prior studies that relied on standalone methods, this research evaluates the comparative advantages of these techniques over alternative clustering and classification approaches. The study also explores emerging trends such as real-time personalization and ethical challenges related to AI-driven targeting.
This article explores the utilization of the Hadoop ecosystem as a polyglot big data processing platform, focusing on the integration of diverse computation and storage technologies and their potential advantages in certain computational contexts. It delves into the potential of this ecosystem as a unified platform highlighting its architectural foundations and their complementary strengths in distributed storage, processing efficiency and real-time analytics. The article explores potential use cases within domains such as Smart Cities and Social Networks, illustrating how the platform's diverse components can be orchestrated in a polyglot manner and how these fields can benefit from the ecosystem's capabilities. Finally, the article concludes by showcasing alternatives for future research, including specialized architectural aspects of the ecosystem to advance the polyglot paradigm.
Machine learning has been essential in enhancing the results of skill acquisition in online learning education, which has seen tremendous growth. This review of the literature focuses on studies that attempted to develop certain competencies via online education by means of machine learning. The integration of machine learning into online learning environments has introduced transformative opportunities to personalize and enhance the educational experience for diverse learners. Online learning encompasses various techniques, such as online supervised, unsupervised, and limited feedback learning, which adapt to data streams and provide scalable solutions for real-time model updates. These capabilities offer significant advantages, including efficient learning tailored to individual needs, improved engagement, and adaptability in dynamic educational contexts. This paper explores the methodologies of online learning and the impact of machine learning on personalizing online education. Key approaches to personalization include adaptive content delivery, real-time performance feedback, and AI-driven support systems such as chatbots, which facilitate continuous engagement and foster self-regulated learning. Institutions can better react to interruptions and assist distant learners using AI-powered adaptive learning, which has been highlighted by the COVID-19 pandemic. As the demand for flexible and accessible learning solutions grows, machine learning stands as a vital tool in advancing personalized online education.
Over the past few years, wildfires have become a worldwide environmental emergency, resulting in substantial harm to natural habitats and playing a part in the acceleration of climate change. Wildfire management methods involve prevention, response, and recovery efforts. Despite improvements in detection techniques, the rising occurrence of wildfires demands creative solutions for prompt identification and effective control. This research investigates proactive methods for detecting and handling wildfires in the United States, utilizing Artificial Intelligence (AI), Machine Learning (ML), and 5G technology. The specific objective of this research covers proactive detection and prevention of wildfires using advanced technology; Active monitoring and mapping with remote sensing and signaling leveraging on 5G technology; and Advanced response mechanisms to wildfire using drones and IOT devices. This study was based on secondary data collected from government databases and analyzed using descriptive statistics. In addition, past publications were reviewed through content analysis, and narrative synthesis was used to present the observations from various studies. The results showed that developing new technology presents an opportunity to detect and manage wildfires proactively. Utilizing advanced technology could save lives and prevent significant economic losses caused by wildfires. Various methods, such as AI-enabled remote sensing and 5G-based active monitoring, can enhance proactive wildfire detection and management. In addition, super intelligent drones and IOT devices can be used for safer responses to wildfires. This forms the core of the recommendation to the fire Management Agencies and the government.
Action Rules are rule based systems that extract actionable patterns which are hidden in big volumes of data. Users need recommendations on actions they can undertake to increase their profit or accomplish their goals, this recommendations are provided by Actionable patterns. In the technological world of big data, massive amounts of data are collected by organizations, including in major domains like financial, medical, social media and Internet of Things(IoT). To analyze and store such a massive amount of data, distributed computing frameworks like Hadoop and Spark are introduced to store the big data in a distributed fashion which manage and analyze them in parallel. The traditional Action Rules extraction models, which analyze the data in a nondistributed fashion, do not perform well when dealing larger datasets. Serious complications of discovering Action Rules with such distributed environments are - data distribution among computing nodes and calculation of major parameters including : support, confidence, utility, and coverage, that represent the whole data. Information granules form basic entities in the world of Granular Computing(GrC), which represents meaningful smaller units derived from a larger complex information system. In this research, we focus on the data distribution phase of the distributed Actionable Pattern Mining problem. To handle the data distribution task by splitting the big data in both horizontal and vertical fashions - we propose partition threshold rho. In this work, we concentrate on using information granules to implement a vertical data splitting strategy with Meta Actions. Hence our results discover valuable Actionable Knowledge with application in Business and Education domains.
Any organization engaged in trading aims to maximize earnings while maintaining costs at their bare minimum. One of the inexpensive ways to accomplish this objective is through sales forecasting.Evidence from empirical literature has shown that sales forecasting frequently results in better customer service, fewer returns of goods, less dead stock, and effective production scheduling. Successful sales forecasting systems are essential for the food sector because of the limited shelf life of food goods and thesignificance of product quality. In this paper, we generated sales of forecasts for a perishable dairy drink using the famous ARIMA approach. We identified the ARIMA (0, 1, 1)(0, 1, 1)12 as the proper model formodeling the daily sales forecast of the perishable drink. After performing model diagnostics, the modelsatisfied all the model assumptions, and a strong positive linear relationship (R2 > 0.9) was observed when the actual daily sales were regressed against the forecasted values.
Depression is a prevailing mental disturbance affecting an individual’s thinking and mental development. There has been much research demonstrating effective automated prediction and detection of Depression.Many datasets used suffer from class imbalance where samples of a dominant class outnumber the minority class that is to be detected. This review paper uses the PRISMA review methodology to enlist different class imbalance handling techniques used in Depression prediction and detection research. The articles were taken from information technology databases. The research gap was found that under sampling methods were few for predicting and detecting Depression and regression modelling could be considered for future research. The results also revealed that the common data level technique is SMOTE as a single method and the common ensemble method is SMOTE, oversampling and under sampling techniques. The model level consisted of various algorithms that can be used to tackle the class imbalance problem.