
This paper provides a comprehensive analysis of architectural patterns and optimization techniques for time-series data processing, centered on the replacement of native DATE/TIMESTAMP types with integerbased surrogate keys. We demonstrate that employing 32-bit (YYYYMMDD) and 64-bit integer formats for datetime representation, coupled with specialized algorithms for indexing, range search, and aggregation, yields substantial performance gains. Empirical evaluations confirm a 30–60% reduction in storage footprint, a 25–40% acceleration in query execution, and up to an eightfold increase in system throughput through batched operations. Beyond these metrics, the study delves into advanced practical implementations across high-frequency trading, telecommunications, and industrial IoT, detailing extended use cases such as real-time fraud detection and predictive maintenance. The paper further introduces a set of actionable implementation guidelines, including hybrid data models and optimized partitioning strategies, to facilitate adoption. Finally, we explore the application of International Atomic Time (TAI) to eliminate temporal ambiguities and outline future research directions integrating this approach with machine learning pipelines and edge computing architectures. The collective findings position integer-based timestamp storage as a foundational element in the design of high-performance, scalable, and reliable time-series data warehouses.
In this study, we optimize SQL+ML queries on top of OpenMLDB, an open-source database that seamlessly integrates offline and online feature computations. The work used feature-rich synthetic dataset experiments in Docker, which acted like production environments that processed 100 to 500 records per batch and 6 to 12 requests per batch in parallel. Efforts have been concentrated in the areas of better query plans, cached execution plans, parallel processing, and resource management. The experimental results show that OpenMLDB can support approximately 12,500 QPS with less than 1 ms latency, outperforming SparkSQL and ClickHouse by a factor of 23 and PostgreSQL and MySQL by 3.57 times. This study assessed the impact of optimization and showed that query plan optimization accounted for 35
Enterprises pursuing AI-driven transformation face a critical tradeoff: centralized consistency vs. decentralized scalability. The "Data Platform Unification Paradox" captures this dilemma. Building on our prior NLPI 2025 paper, this extended version integrates technical depth, mathematical models, and concrete architectures, especially for integrating Data Mesh with Quantum Databases and LLM Agents. A federated architecture is proposed using graph-theoretic models and entropy-based data valuation. We introduce a formal structure to evaluate platform complexity and propose intelligent agent-based governance models to operationalize data sharing across domains. This work aims to move beyond conceptual frameworks by proposing actionable blueprints for next-generation, intelligent data ecosystems.
Agile methodologies have transformed organizational management by prioritizing team autonomy and iterative learning cycles. However, these approaches often lack structured mechanisms for knowledge retention and interoperability, leading to fragmented decision-making, information silos, and strategic misalignment. This study proposes an alternative approach to knowledge management in Agile environments by integrating Ikujiro Nonaka and Hirotaka Takeuchi’s theory of knowledge creation— specifically the concept of Ba, a shared space where knowledge is created and validated—with Jürgen Habermas’s Theory of Communicative Action, which emphasizes deliberation as the foundation for trust and legitimacy in organizational decision-making. To operationalize this integration, we propose the Deliberative Permeability Metric (DPM), a diagnostic tool that evaluates knowledge flow and the deliberative foundation of organizational decisions, and the Communicative Rationality Cycle (CRC), a structured feedback model that extends the DPM, ensuring long-term adaptability and data governance. This model was applied at Livelo, a Brazilian loyalty program company, demonstrating that structured deliberation improves operational efficiency and reduces knowledge fragmentation. The findings indicate that institutionalizing deliberative processes strengthens knowledge interoperability, fostering a more resilient and adaptive approach to data governance in complex organizations.
Business process intelligence improves operational efficiency that is essential for achieving business objectives, besides facilitating competitive advantage. As organizations operate through inter-connected business processes, insights into their process performance through related business rules are essential to achieve business objectives. This paper outlines an approach for developing insight-driven business rules using the concept of buckets, which can lead to the development of a repository of business knowledge for business process operations. The proposed concepts are demonstrated through a prototype modeled on a hypothetical customer mortgage lending process, implemented using Oracle’s PL/SQL database language.
With the rapid growth of Internet finance, competition within the banking industry has intensified significantly. To better understand customer needs and enhance customer loyalty, it has become crucial to develop a customer churn prediction model. Such a model enables banks to identify customers at risk of leaving, support data-driven business decisions, and implement strategies to retain valuable clients, thereby safeguarding the bank's interests. In this context, this paper presents a customer churn prediction model based on an ensemble learning algorithm. Experimental results demonstrate that the model effectively predicts and analyzes potential customer churn, providing valuable insights for retention efforts.
Data anonymization is one of the solutions allowing companies to comply with the GDPR directive in terms of data protection. In this context, developers must follow several steps in the process of data anonymization in development and testing environments. Indeed, real personal and sensitive data must not leave the production environment which is very secure. Often, anonymization experts are faced with difficulties including the lack of data flows and mapping between data sources, the non-cooperation of the database project teams (refusal to change) or even the lack of skills of these teams present due to the age of the systems developed by experienced teams who unfortunately left the project. Other problems are lack of data models. The aim of this paper is to discuss an anonymization process of databases of banking applications and present our context-based recommendations to overcome the different issues met and the solutions to improve methodologies of data anonymization process.
Educational research often encounters clustered data sets, where observations are organized into multilevel units, consisting of lower-level units (individuals) nested within higher-level units (clusters). However, many studies in education utilize tree-based methods like Random Forest without considering the hierarchical structure of the data sets. Neglecting the clustered data structure can result in biased or inaccurate results. To address this issue, this study aimed to conduct a comprehensive survey of three tree- based data mining algorithms and hierarchical linear modeling (HLM). The study utilized the Programme for International Student Assessment (PISA) 2018 data to compare different methods, including non-mixed- effects tree models (e.g., Random Forest) and mixed-effects tree models (e.g., random effects expectation minimization recursive partitioning method, mixed-effects Random Forest), as well as the HLM approach. Based on the findings of this study, mixed-effects Random Forest demonstrated the highest prediction accuracy, while the random effects expectation minimization recursive partitioning method had the lowest prediction accuracy. However, it is important to note that tree-based methods limit deep interpretation of the results. Therefore, further analysis is needed to gain a more comprehensive understanding. In comparison, the HLM approach retains its value in terms of interpretability. Overall, this study offers valuable insights for selecting and utilizing suitable methods when analyzing clustered educational datasets.
This research advocates support for the continuous development and modernization of the police and gendarmic forces, analysing the use of Drones in the operational activities of the Portuguese and Spanish security, police and gendarmic forces: the GNR and the Guardia Civil. Analysing the implementation and expansion of Drones, valuing how the use of these means is advantageous for the police service and for the operations of the GNR and Guardia Civil. Due to the major changes taking place in the world, it is crucial to rethink security and Portugal is gradually adapting to this reality, resulting in new demands for daily police service. The adopted methodology is based on the inductive method that allowed the data collected through analysis to be generalized of data on Drones of the Guardia Civil, appreciating their characteristics and use, with the aim of understanding and comparing their modus operandi regarding the use of Drones in the GNR. In short, it was possible to verify the importance and potential of Drones in surveillance, reconnaissance and target tracking missions having been carried out productive and important conclusions for building Drones capacity in the GNR.
R is widely used by researchers in the statistics field and academia. In Botswana, it is used in a few research for data analysis. The paper aims to synthesis research conducted in Botswana that has used R programming for data analysis and to demonstrate to data scientists, the R community in Botswana and internationally the gaps and applications in practice in research work using R in the context of Botswana. The paper followed the PRISMA methodology and the articles were taken from information technology databases. The findings show that research conducted in Botswana that use R programming were used in Health Care, Climatology, Conservation and Physical Geography, with R part as the most used R package across the research areas. It was also found that a lot of R packages are used in Health care for genomics, plotting, networking and classification was the common model used across research areas.
To process a large volume of data, modern data management systems use a collection of machines connected through a network. This paper proposes frameworks and algorithms for processing distributed joins—a compute- and communication-intensive workload in modern data-intensive systems. By exploiting multiple processing cores within the individual machines, we implement a system to process database joins that parallelizes computation within each node, pipelines the computation with communication, parallelizes the communication by allowing multiple simultaneous data transfers (send/receive). Our experimental results show that using only four threads per node the framework achieves a 3.5x gains in intra-node performance while compared with a single-threaded counterpart. Moreover, with the join processing workload the cluster-wide performance (and speedup) is observed to be dictated by the intra-node computational loads; this property brings a near-linear speedup with increasing nodes in the system, a feature much desired in modern large-scale data processing system.
Deep learning has been well used in many fields. However, there is a large amount of data when training neural networks, which makes many deep learning frameworks appear to serve deep learning practitioners, providing services that are more convenient to use and perform better. MindSpore and PyTorch are both deep learning frameworks. MindSpore is owned by HUAWEI, while PyTorch is owned by Facebook. Some people think that HUAWEI's MindSpore has better performance than FaceBook's PyTorch, which makes deep learning practitioners confused about the choice between the two. In this paper, we perform analytical and experimental analysis to reveal the comparison of training speed of MIndSpore and PyTorch on a single GPU. To ensure that our survey is as comprehensive as possible, we carefully selected neural networks in 2 main domains, which cover computer vision and natural language processing (NLP). The contribution of this work is twofold. First, we conduct detailed benchmarking experiments on MindSpore and PyTorch to analyze the reasons for their performance differences. This work provides guidance for end users to choose between these two frameworks.
Clustering is a crucial part in the field of data mining, and common clustering methods include division-based methods, hierarchy-based methods, density-based methods, and grid-based methods. In order to improve the accuracy of clustering, an optimization study is made mainly for the division-based method FCM clustering, and an FCM clustering method that integrates active learning and principal component analysis (PCA) is proposed. The method first uses principal component analysis to reduce the dimensionality of the data to reduce the computation of electricity data, then trains the sample model by active learning, and introduces the entropy (Entropy) method in the uncertainty sampling method, the larger the entropy means the greater the uncertainty of the sample, and the smaller the entropy means the smaller the uncertainty of the sample, so as to filter the electricity data, and finally the electricity data are clustered by FCM clustering The power data is finally categorized by FCM clustering, and with the proliferation of power data, the power data can be more accurately categorized using this method to achieve the stability of the power grid as well as the utilization rate. Experimental results on three datasets show that this method improves the accuracy of power data clustering by up to 2 percentage points compared to the traditional clustering method without active learning, and achieves good results in each dataset compared to other methods.
The growth of big-data sectors such as the Internet of Things (IoT) generates enormous volumes of data. As IoT devices generate a vast volume of time-series data, the Time Series Database (TSDB) popularity has grown alongside the rise of IoT. Time series databases are developed to manage and analyze huge amounts of time series data. However, it is not easy to choose the best one from them. The most popular benchmarks compare the performance of different databases to each other but use random or synthetic data that applies to only one domain. As a result, these benchmarks may not always accurately represent real-world performance. It is required to comprehensively compare the performance of time series databases with real datasets. The experiment shows significant performance differences for data injection time and query execution time when comparing real and synthetic datasets. The results are reported and analyzed.
Polycystic ovarian syndrome(PCOS) is one of the predominant hormonal imbalances present in women of reproductive age. It needs to be diagnosed and treated at an earlier stage as it's inter-related to diabetes, high cholesterol levels, and obesity. This paper presents an application specially designed for women to help them keep track of their Body Mass Index, Blood Sugar, and Blood Pressure based on their age. The people diagnosed with PCOS(an endocrine disorder) can use this application to make their life easy since it helps follow certain exercises, diets, and timely reminders for water and medicines. It has features like the period tracker to track the user’s menstrual cycle, find dieticians nearby, links to various PCOS supplements, users can track their moods during different menstrual phases and control their mood swings. Finally, the application has games to add that interactive touch.
The integration of XML data sources which have different schemas/DTD can originate structural and vocabular heterogeneity. In this context, it is difficult to write satisfiable queries. As a solution, many Information Systems focus on building approximate evaluation techniques for exact queries. As a project, we build flexible and preference XML query languages and associated evaluation algorithms. In this paper, we propose the Flexible Preference Tree Pattern Query (FPTPQ), a new TPQ that allows multiple items/names (resp. paths) for the same node, in order to integrate (resp. to locate) all the different instances of the database nodes. The FPTPQ enable to have preference nodes and ordering operators among label items and paths. We also provide a holistic algorithm that evaluates the FPTPQ and capitalises the preferences to determine the best available solutions. Illustrations and experimentations are realized to show the effectiveness of our solutions.
In cloud environments, hardware configurations, data usage, and workload allocations are continuously changing. These changes make it difficult for the query optimizer of a cloud database management system (DBMS) to select an optimal query execution plan (QEP). In order to optimize a query with a more accurate cost estimation, performing query re-optimizations during the query execution has been proposed in the literature. However, some of there-optimizations may not provide any performance gain in terms of query response time or monetary costs, which are the two optimization objectives for cloud databases, and may also have negative impacts on the performance due to their overheads. This raises the question of how to determine when are-optimization is beneficial. In this paper, we present a technique called ReOptML that uses machine learning to enable effective re-optimizations. This technique executes a query in stages, employs a machine learning model to predict whether a query re-optimization is beneficial after a stage is executed, and invokes the query optimizer to perform the re-optimization automatically. The experiments comparing ReOptML with existing query re-optimization algorithms show that ReOptML improves query response time from 13% to 35% for skew data and from 13% to 21% for uniform data, and improves monetary cost paid to cloud service providers from 17% to 35% on skewdata.
Oracle is one of the largest vendors and the best DBMS solution of Object Relational DBMS in the IT world. Oracle Database is one of the three market-leading database technologies, along with Microsoft SQL Server's Database and IBM's DB2. Hence in this paper, we have tried to answer the million-dollar question “What is user’s responsibility to harden the oracle database for its security?” This paper gives practical guidelines for hardening the oracle database, so that attacker will be prevented to get access into the database. The practical lookout for protecting TNS, Accessing Remote Server and Prevention, Accessing Files on Remote Server, Fetching Environment Variables, Privileges and Authorizations, Access Control, writing security policy, Database Encryption, Oracle Data Mask, Standard built in Auditing and Fine Grained Auditing (FGA) is illustrated with SQL syntax and executed with suitable real life examples and its output is tested and verified. This structured method acts as Data Invictus wall for the attacker and protect user’s database.
In this paper, we describe the developed model of the Convolutional Neural Networks CNN to a classification of advertisements. The developed method has been tested on both texts (Arabic and Slovak texts).The advertisements are chosen on a classified advertisements websites as short texts. We evolved a modified model of the CNN, we have implemented it and developed next modifications. We studied their influence on the performing activity of the proposed network. The result is a functional model of the network and its implementation in Java and Python. And analysis of model results using different parameters for the network and input data. The results on experiments data show that the developed model of CNN is useful in the domains of Arabic and Slovak short texts, mainly for some classification of advertisements. This paper gives complete guidelines for authors submitting papers for the AIRCC Journals.
Distributed databases and data replication are effective ways to increase the accessibility and reliability of un-structured, semi-structured and structured data to extract new knowledge.Replications offer better performance and greater availability of data.With the advent of Big Data, new storage and processing challenges are emerging.To meet these challenges, Hadoop and DHTs compete in the storage domain and MapReduce and others in distributed processing, with their strengths and weaknesses.We propose an analysis of the circular and radial replication mechanisms of the CLOAK DHT.We evaluate their performance through a comparative study of data from simulations.The results show that radial replication is better in storage, unlike circular replication, which gives better search results.