2025 27th International Conference on Business Informatics (CBI)(2025)
German Research Center for Artificial Intelligence
被引用0|浏览0
摘要
As data-driven applications gain traction in smart living environments, ensuring high data quality becomes a critical prerequisite for reliable analytics and Artificial Intelligence (AI)based services. However, due to the wide range of hardware and software providers as well as the variety of use cases in the smart living domain, the resulting data are characterized by highly heterogeneous structures and nesting levels. The diversity of systems, their strong reliance on context, and the rapid growth of Internet of Things (IoT) devices make it difficult to ensure consistency, accuracy, and reliability in data quality assessments within smart living environments. This paper addresses the need for a consistent assessment of data quality within heterogeneous smart living datasets. We present a prototype that generates quality reports for datasets with an unknown structure based on established data quality metrics A Large Language Model component generates structured metadata describing the dataset's structure and suggesting applicable quality checks, which then serve as input for an implemented code base to perform standardized quality checks. This approach not only supports the assessment of data quality within individual organizations but also facilitates cross-organizational assessments, enabling better comparability and evaluation of datasets from different sources or providers. The prototype was evaluated using a variety of synthetic and real-world smart living datasets. The results demonstrate the feasibility of the proposed approach, although certain limitations remain, which are discussed in detail within the paper.
更多
查看译文
关键词
Data Quality Assessment,Smart Living,Data Structures,Heterogeneous Datasets