Common elementary price indices include Dutot, Carli, and Jevons, while less-known ones include Carruthers-Sellwood-Ward-Dal & eacute;n (CSWD) and harmonic indices. Recently, a new elementary index, the Dikhanov index, has been proposed, previously introduced independently by Allyn Young, Bert Balk, and Jens Mehrhoff. This paper presents the axiomatic properties of the Young-Balk-Mehrhoff-Dikhanov (YBMD) index and compares it with population elementary indices under log-normal price assumptions. Using simulations, biases, and Mean Squared Errors (MSEs) of sample indices are analyzed. Results show that the YBMD index is asymptotically unbiased, with bias and MSE decreasing as price correlations between periods increase. Notably, the study identifies scenarios where the sample YBMD index exhibits lower bias and MSE than the sample CSWD and Jevons indices, highlighting its potential advantages in practical applications.
The procedure used by a National Statistical Office (NSO) for collecting prices to produce the Consumer Price Index (CPI) is based on sample surveys. The universe (or population) of items has three dimensions: product, geographical, and time, all of which are described in the paper. This paper presents and discusses general concepts and techniques of survey sampling that are crucial for the construction of price indices. In particular, both probability and non-probability sampling techniques are discussed and illustrated with the real-world examples. A separate section discusses sampling scanned products. One of the approaches used for such data is the dynamic approach, which involves monthly sampling by applying appropriate data filters. This technique can be seen as a special form of cut-off s ampling. The empirical study investigates the effect of data filtering on the level of price indices. The main pragmatic conclusion is that the low-sales filter has the most significant impact on reducing the size of the scanner dataset. The second important conclusion is that changing the order of data filtering has minimal impact on the value of the price index.
Price index decomposition allows the National Statistical Institute to break down aggregated price movements into contributions of individual commodities or product groups. In particular, decomposing multilateral indices, which combine many different time comparisons within the time window, can be useful in understanding and interpreting them by common users. This paper discusses multiplicative decompositions of multilateral GEKS-type indices, among which, currently, the GEKS-F and GEKS-T formulas are the most widely used. The paper also provides a multiplicative decomposition of the known GEKS-W index, as well as a decomposition of the GEKS-L, GEKS-GL, and GEKS-LM multilateral indices recently proposed in the literature. The paper also proposes normalized multiplicative decompositions of multilateral indices along with a relative commodity impact measure to enable comparisons of commodity contributions between products and across indices. The effects of these decompositions are demonstrated on two real scanner data sets from the food product segment.
Scanner data from retailers like supermarkets, electronics stores, and online shops provides detailed transaction information at the barcode level (e.g., GTIN, EAN), allowing for the use of various price index formulas, including weighted ones. Due to high product turnover and seasonality, multilateral index methods are ideal, as they use a whole-time window and are transitive, avoiding chain drift. However, most commonly used multilateral indices (e.g., GEKS, CCDI, GK, TPD) fail the identity test, which requires the index to return to one when prices revert to their original levels. This paper proposes a new multilateral index inspired by GEKS but incorporating quality adjustments like the Geary-Khamis method. The index satisfies the identity test and other key axioms, demonstrating its robustness. Comparisons with the SPQ index and quality-adjusted indices (GEKS-AQU, GEKS-AQI) confirm its effectiveness, making it a highly useful tool for scanner data analysis in both theory and practice.
This article addresses the problem of selecting a price index dedicated to scanner data, with attention directed toward multilateral GEKS-type indices. The article proposes three general classes of GEKS-type indices, the special cases of which are reduced to the well-known GEKS, CCDI, or GEKS-W formulas. The main purpose of the article is to analyze the axiomatic properties of the proposed index classes. The practical conclusion of the analysis seems to be the recommendation of a general class of GEKS-type indices, which is based on the Lloyd-Moulton index and of which the recently published GEKS-L and GEKS-GL indices are special cases.
An important challenge for official statistics is the automatic detection of downsized and upsized products generating the price index bias. This paper presents the results of the authors' research on the phenomenon of downsizing and upsizing with the use of scanner data. The contribution of the paper is as follows: (1) the paper systematizes, defines and delineates all the potentially dangerous situations for measuring inflation associated with a change in product size; (2) the paper discusses the potential problems that the statistical office may encounter when trying to detect downsized and upsized products; (3) the paper proposes an automatic procedure for selecting downsized and upsized products on the basis of selected food and non-food groups of scanned products; (4) the paper presents the results of empirical studies that clearly show the impact of downsizing and upsizing on the bias of inflation measurement, with the impact assessed from the perspective of commonly known bilateral and multilateral price indices.
Although modern price index theory is based on comparisons of ratios of prices, quantities and expenditures, we may be more interested in the magnitude of differences in these characteristics in many business applications. The benefit of using these differences is that there is no problem associated with the occurrence of zero prices and quantities, a problem that arises when we work with ratios. In practice, we most often care about decomposing value difference into indicators of contributions from price and quantity differences. The best-known price and quantity indicators are the Bennet indicators, which are not transitive. Although there have been papers in the literature that propose a transitive version of the Bennet indicators, they deal with comparisons across firms in cross-section or panel contexts. This paper revises the price and quantity Bennet indicators and their multilateral versions for the analysis of scanner data. Specifically, i nstead o f c onsidering c omparisons across firms, c ountries or r egions, the t ransitive versions of t he Bennet i ndicators a re a dapted to work on scanner data sets observed over a fixed time w indow. Since the scanner data sets have a high turnover of products, which can make it difficult to interpret the difference in sales values over the compared time periods, the paper also considers a matched sample approach. One of the objectives of the study is to compare bilateral and multilateral Bennet indicator results across all available products or strictly matched products over time. It also examines the impact of data filters used and the level of data aggregation on the price and quantity Bennet indicators. According to the best author’s knowledge, this study is a pioneer in the field of implementing the multilateral Bennet indicators in scanner data analysis.
The aim of the article is to examine the major aspects of Poland's economic situation as reflected in the main macroeconomic indicators during the country’s twenty-year membership in the European Union (EU). The analysis is based on reports prepared by the International Monetary Fund (IMF), the National Bank of Poland (NBP), the Organisation for Economic Co-operation and Development (OECD), and statistical data provided by Eurostat and the Central Statistical Office of Poland (Statistics Poland). The article discusses issues related to the growth of Gross Domestic Product (GDP) and the improvement in the standard of living of Polish citizens, which were significantly stimulated by Poland's accession to the EU. Membership in the European Union brought Poland many tangible economic benefits. It became a key factor stimulating economic development through access to the single European market and EU structural funds, supporting investment, development, and infrastructure improvement. EU membership opened new trading opportunities for Poland, facilitating the export of goods and services to European markets. EU membership also enabled the free movement of citizens between the Member States. These benefits have helped modernise the country and enhance Poland's position on the international stage.
Scanner data are electronic transaction data that specify turnover and the number of items sold by barcodes, e.g., the Global Trade Article Number. These data are of particular value and interest to theorists and practitioners who wish to measure the Cost of Living Index or the Consumer Price Index, since their complete content makes it possible to compute any price index formula, including superlative indices or CES (Constant Elasticity Substitution) indices. Since the CES index requires the estimation of the elasticity of substitution, this paper focuses on verifying various methods of estimating this parameter based on scanner data. The paper considers both algebraic methods and methods based on the panel regression approach. The main achievement of the paper is the separation of the main factors that affect the estimated value of the elasticity of substitution, i.e., the type of data filter used and the level of data aggregation. The paper also verifies how the elasticity of substitution estimates affect the differences between the values of the CES indices based on these estimates.
A wide variety of retailers (supermarkets, home electronics, Internet shops, etc.) provide scanner data containing information at the level of the barcode, e.g. the Global Trade Item Number (GTIN). As scanner data provide complete transaction information, we may use the expenditure shares of items as weightsfor calculating price indices at the lowest (elementary) level of data aggregation. The challenge here is the choice of the index formula which should be able to reduce chain drift bias and substitution bias. Multilateral index methods seem to be the best choice due to the dynamic character of scanner data. These indices work on a wholetime window and are transitive, which is key to the elimination of the chain drift effect. Following what is called an identity test, however, it may be expected that even when only prices return to their original values, the index becomes one. Unfortunately, the commonly used multilateral indices (GEKS, CCDI, GK, TPD, TDH) do not meet the identity test. The paper discusses the proposal of two multilateral indices and their weighted versions. On the one hand, the design of the proposed indices is based on the idea of the GEKS index. On the other hand, similarly to the Geary-Khamis method, it requires quality adjusting. It is shown that the proposed indices meet the identity test and most other tests. In an empirical and simulation study, these indices are compared with the SPQ index, which is relatively new and also meets the identity test. The analytical considerations as well as empirical studies confirm the high usefulness of the proposed indices.
Scanner data mean electronic transaction data that specify product prices and their expenditures obtained from supermarkets’ IT systems by scanning bar codes (i.e. GTIN or SKU). Scanner data are a relatively new and cheap data source for the calculation of the Consumer Price Index (CPI) and the biggest advantage of scanner data is the full product information they provide already at the lowest level of aggregation. Thus, the digitization of the public sector becomes not only something that is needed but an actual necessity resulting from organisational and economic premises (e.g.: reduction of costs or time related to obtaining data). One of main challenges while using scanner data is the choice of the right price index. The list of potential price indices, which could be used in the scanner data case, is quite wide, i.e. bilateral and multilateral indices are used in practice. One of the most important criterion in selecting index formula for scanner data case is the potential reduction of the chain drift bias. The chain drift occurs if the index differs from unity when prices and quantities revert back to their base level. In the paper we present situations on the market leading to the serious chain drift bias. Our main hypothesis is that lagging consumers’ reaction to price changes is the cause of the chain drift effect. Moreover, the article is an attempt to answer the question whether the correlation of prices and quantities may have an influence on the scale and sign of the bias of the measurement of price dynamics. The study focuses also on the scale of over- and underestimation the target full-window multilateral indices by their corresponding splicing extensions. Finally, the paper verifies a hypothesis that the identity test is a key property in reducing chain drift bias. In order to verify the above research problems, both empirical and simulation studies were carried out. Our main result is the confirmation of earlier suspicions that delayed consumer response and price-quantity correlation are determinants of chain drift bias.
Scanner data are electronic transaction data most often from retail chains and obtained from electronic retail terminals. The identification of products takes place after scanning their characteristic barcode (e.g. EAN or GTIN), thus in the case of scanner data, we have full product information (price, sales volume, weight, description, etc.) at the most disaggregated level. In the cases of many countries, as well as Poland, this type of data is a valuable alternative source of information when estimating inflation. This paper discusses the main advantages but also the challenges of using scanner data in the CPI measurement. The main purpose of the paper, however, is to discuss the problem of selecting an optimal price index formula that would be appropriate for the highly dynamic (in terms of product rotation) scanner data. The considerations, supported by examples of empirical studies, will be demonstrated using the PriceIndices package in the R environment.
Scanner data can be obtained from a wide variety of retailers (supermarkets, home electronics, Internet shops, etc.) and provide information at the level of the barcode, i.e. the Global Trade Item Number (GTIN, formerly known as the EAN code). After cleaning data and unifying product names, products should be carefully classified (e.g. into the COICOP 5 level or below), matched, filtered, and aggregated. These procedures often require creating new IT or writing custom scripts (R, Python, Mathematica, SAS, others). One of new challenges connected with scanner data is the appropriate choice of the index formula. The article discusses a new R package, i.e. PriceIndices, which is used to process scanner data and to calculate bilateral and multilateral price indices, along with their window extensions. The assumptions for the construction of the package were such that it would serve both practitioners and scientists through a multitude of methods and their parametrization. The main purpose of the article is to present the utility of the package in the field of analyzing the dynamics of scanner prices.
Abstract One of the greatest challenges facing official statistics in the 21st century is the use of alternative sources of data about prices (scanned and scraped data) in the analysis of price dynamics, which also involves selecting the appropriate formula of the price index at the elementary group (5-digit) level. When consumer price indices of goods and services are constructed, a number of subjective decisions are made at different stages, e.g. regarding the choice of data sources and types of indices used for the purpose of estimation. All of these decisions can affect the bias of consumer price indices, i.e. the extent to which they contribute to the overall uncertainty about the resulting index values. By measuring how robust consumer price indices are, one can assess the impact that the decisions made at the different stages of index construction have on the index values. This assessment involves analysing uncertainty and sensitivity. The purpose of the study described in the article was to determine how much and in which direction the consumer price index changes when including scanner and scraped data in the analysis, in addition to the data on prices collected by enumerators. The impact of these new data sources was assessed by analysing uncertainty and sensitivity under the deterministic approach. To the best of the authors’ knowledge, it is a novel application of robustness analysis to measure inflation using new data sources. The empirical study was based on data for February and March 2021, while scanner and scraped data about selected categories of food products were obtained from one retail chain operating hundreds of points of sale in Poland and selling products online. It was found that the choice of a data source has the most significant impact on the final value of the index at the elementary group level, while the choice of the aggregation formula used to consolidate different data sources is of secondary importance.
The article deals with the problem of the proper selection of the theoretical distribution to describe the empirical distribution of scanner prices. In the empirical study we use scanner data from one retail chain in Poland, i.e. monthly data on natural yoghurt, drinking yoghurt, long grain rice and coffee powder sold in 212 outlets in January and February 2022. Prices and price relatives were modeled using selected ten probability distributions with non-negative support, including two, three and four-parameter family of distributions In addition to the visual assessment in the form of empirical PDF and CDF figures, numerical criteria were used. These include information criteria values such as AIC, BIC, HQIC and p values calculated for the K-S, AD and CVM goodness-of-fit tests. Our research showed that at least two models could be distinguished as very accurate, which provides a good background for simulation research on price indices or for the construction of so-called population price indices.
Scanner data can be obtained from a wide variety of retailers (supermarkets, home electronics, Internet shops, etc.) and provide information at the level of the barcode, i.e. the Global Trade Item Number or its European version: European Article Number. One of advantages of using scanner data in the Consumer Price Index measurement is the fact that they contain complete transaction information, i.e. prices and quantities for every sold item. One of new challenges connected with scanner data is the choice of the index formula which should be able to reduce the chain drift bias and the substitution bias. Multilateral index methods seem to be the best choice in the case of dynamic scanner data sets. These indices work on a whole time window and are transitive, which is a key property in eliminating the chain drift effect. Following the so-called identity test, however, one may expect that even when only prices return to their original values, the index becomes one. Unfortunately, the commonly used multilateral indices (GEKS, CCDI, GK, TPD, TDH) do not meet the identity test. The paper discusses the proposal of two multilateral indices, the idea of which resembles the GEKS index, but which meet the identity test and most of other tests. In an empirical study, these indices are compared, inter alia, with the SPQ index, which is relatively new and also meets the identity test. Analytical considerations as well as empirical study confirm the high usefulness of the proposed indices.
Scanner data mean transaction data that specify product prices and their expenditures obtained from supermarkets' IT systems by scanning bar codes (i.e. GTIN or SKU). Scanner data are a relatively new and cheap data source for the calculation of the Consumer Price Index (CPI) and the main advantage of using scanner data is the fact that they provide full information about products even on the lowest data aggregation level. One of main challenges while using scanner data is the choice of the appropriate price index formula. The list of potential price indices, which could be used in the scanner data case, is quite wide, i.e. bilateral and multilateral indices are used in practice. For instance, some countries use the chain Jevons price index formula while some other countries prefer the multilateral GEKS index or the Geary-Khamis method. One of the most important criterions in selecting index formula for scanner data case is the potential reduction of the chain drift bias. The chain drift occurs if the index differs from unity when prices revert back to their base level. In the paper we present some simulation results which show the situations on the market leading to the serious chain drift bias.
The methodology of price indices dedicated to scanner data is broad, multifaceted, and still contains many open problems. The main challenges include choosing the index formula and the time window width in multilateral methods, as well as determining splicing or other data updating methods. Many NSIs experiment with scanner data, their processing, classifying, matching, and finally using this type of data for CPI calculations. However, these activities are limited, which is partly due to the lack of widely available software in this field. On the one hand, R packages dedicated to price indices are available (e.g. IndexNumR or micEconIndex), on the other hand, their functionality and the scope of implemented methods are quite limited. The article discusses a new R package, i.e. PriceIndices, which is used to process scanner data and to calculate bilateral and multilateral price indices. The assumptions for the construction of the package were such that it would serve both practitioners and scientists through a multitude of methods and their parameterization. The main purpose of the article is to present the utility of the package in the field of analyzing the dynamics of scanner prices. All obtained results are based on the real scanner data set on milk obtained from one retailer chain in Poland and included in the PriceIndices package.
Scanner data offer new opportunities for CPI or HICP calculation. They can be obtained from a wide variety of retailers (supermarkets, home electronics, Internet shops, etc.) and provide information at the level of the barcode. One of advantages of using scanner data is the fact that they contain complete transaction information, i.e. prices and quantities for every sold item. After clearing data and unifying product names, products should be carefully classified (e.g. into COICOP 5 or below), matched, filtered and aggregated. One of new challenges connected with scanner data is the appropriate choice of the index formula. In this article we present a proposal for the implementation of individual stages of handling scanner data. We also point out potential problems during scanner data processing and their solutions. We compare a large number of price index methods based on real scanner data sets and we verify their sensitivity on adopted data filtering and aggregating methods. One of the aims is also to compare calculations of multilateral indices in terms of how time-consuming they are. Finally, the paper investigates distances between these indices and the theoretical, expected value of the price share when prices are log-normally distributed. It is a new approach to providing an additional criterion in the price index selection.