This study proposes a structural machine learning methodology that integrates both linear and nonlinear relationships for crude oil price forecasting. By employing a partially linear machine learning model that explicitly captures the linear effects of key variables influencing crude oil prices, the approach enhances interpretability while evaluating predictive performance. In addition, this study investigates the impact of hyperparameter selection on forecasting accuracy, with a particular emphasis on subsampling and random seed effects-factors that have received limited attention in existing empirical research. Subsampling is actively utilized as a hyperparameter to explore variations in predictive performance, and instead of relying on a single fixed random seed, multiple seeds are used to assess the model's stability and robustness. Based on monthly forecasting experiments spanning approximately 11 years, the results demonstrate that the partially linear machine learning model, when optimized through appropriate subsampling and hyperparameter tuning, outperforms benchmark models in one- to three-step-ahead forecasts. Furthermore, an analysis of prediction error distributions across different random seeds confirms the robustness of the model's predictive performance.
We employ a semiparametric functional coefficient panel approach to allow an economic relationship of interest to have both country-specific heterogeneity and a common component that may be nonlinear in the covariate and may vary over time. Surfaces of the common component of coefficients and partial derivatives (elasticities) are estimated and then decomposed by functional principal components, and we introduce a bootstrap-based procedure for inference on the loadings of the functional principal components. Applying this approach to national energy-GDP elasticities, we find that elasticities are driven by common components that are distinct across two groups of countries yet have leading functional principal components that share similarities. The groups roughly correspond to OECD and non-OECD countries, but we utilize a novel methodology to regroup countries based on common energy consumption patterns to minimize root mean squared error within groups. The common component of the group containing more developed countries has an additional functional principal component that decreases the elasticity of the wealthiest countries in recent decades.
Despite the fact that China and the United States represent the G2 in terms of economic size, the RMB’s international significance in the existing international financial system is limited. China has made significant progress in encouraging RMB internationalization. It has the ability to disrupt the global financial system, dominated by the US dollar. In order to seize chances under such circumstance, Korea must find a new direction for the internationalization of Korean Won.This collective volume has seven independent papers that investigate the current and future status of the RMB internationalization and its impacts and implications on Korean economy. The summarizations of each paper are as follows: Chapter One explains the background and motivation of this collective volume. Also, it describes the current status of the Korean Won Internationalization. Chapter Two provides detailed descriptions of the current status of RMB internationalization and its future prospects. Chapter Three examines the performance of Shanghai and Seoul RMB-KRW direct foreign exchange markets and figures that such direct FX markets have not been fully developed yet. However, such markets are expected to become more efficient gradually. Chapter Four finds that in the gradual evolution of the RMB internationalization, the KRW is becoming more synchronized with the Chinese yuan. Chapter Five explores the factors for coupling between the RMB and the KRW. Not only trade and finance channels but also policy implementations are important. Chapter Six develops the index for the KRW internationalization in light of the RMB internationalization. This chapter finds that the RMB internationalization may hinder the KRW internationalization. If the Korean government continues to delay the KRW internationalization, the benefit from the currency internationalization will become smaller. In that regard, Chapter Seven emphasizes that at this moment, it is very meaningful to re-examine the long-term strategy for Korea’s won internationalization. The internationalization of the Korean won provides a new opportunity for the country’s financial development, rather than a disruption to the current global financial system. In the process of the internationalization of RMB and KRW, China and Korea should further strengthen bilateral financial and economic cooperation to push forward the process of RMB and KRW internationalization. Because the RMB or the KRW each have such a small proportion of the global monetary system, it is premature to be concerned about competition. Cooperation should take priority. The key policy suggestions for cooperation between these two currency internationalizations could at the very least include: (1) increasing the bilateral currency swap lines (BSLs) and making them more effective; (2) encouraging more usage of RMB and KRW in bilateral trade and direct investment; (3) encouraging more Chinese investors to hold KRW denominated assets and so does the other party; (4) accelerating the development of offshore RMB market in Seoul while having increasingly important offshore KRW market in China; (5) strengthening the coordination and cooperation in exchange rate policies.
The binary indicator of collusion is the key ingredient in estimating overcharges from bid-rigging with a regression-based approach. We develop a method for examining the effects of misclassification error in the indicator of bid-rigging status on estimates of damages from collusion. We derive partial identification of the regression model of winning bids in public procurement auctions and provide informative bounds on the price effects of bid-rigging. We find that the bounds are tight when placing a plausible restriction on the extent of measurement errors. Our findings show that relaxing the nondifferential assumption about misclassification errors leads to wider bounds.
In this paper, we address the issue of the time-varying relationship between health and long-term income and show that income profile over time is more important than the permanent income of a specific period in explaining general health condition. A functional probit regression model is introduced to investigate how the income profiles of middle-aged people can affect the health condition of the last period of the given period using the Panel Study of Income Dynamics data. We also perform the probit estimation to check the robustness of the functional regression results. The empirical results of our paper clearly indicate a time-varying relationship between the entire distribution of income and health status at one point in time, which cannot be accurately explained by the current income or permanent income of the given period.
Previous authors have pointed out that energy consumption changes both over time and nonlinearly with income level. Recent methodological advances using functional coefficients allow panel models to capture these features succinctly. In order to forecast a functional coefficient out-of-sample, we use functional principal components analysis (FPCA), reducing the problem of forecasting a surface to a much easier problem of forecasting a small number of smoothly varying time series. Using a panel of 180 countries with data since 1971, we forecast energy consumption to 2035 for Germany, Italy, the US, Brazil, China, and India.
Current diagnostic standards for lymphoproliferative disorders include multiple tests for detection of clonal immunoglobulin (IG) and/or T-cell receptor (TCR) rearrangements, translocations, copy-number alterations (CNAs), and somatic mutations. The EuroClonality-NGS DNA Capture (EuroClonality-NDC) assay was designed as an integrated tool to characterize these alterations by capturing IGH switch regions along with variable, diversity, and joining genes of all IG and TCR loci in addition to clinically relevant genes for CNA and mutation analysis. Diagnostic performance against standard-of-care clinical testing was assessed in a cohort of 280 B- and T-cell malignancies from 10 European laboratories, including 88 formalin-fixed paraffin-embedded samples and 21 reactive lesions. DNA samples were subjected to the EuroClonality-NDC protocol in 7 EuroClonality-NGS laboratories and analyzed using a bespoke bioinformatic pipeline. The EuroClonality-NDC assay detected B-cell clonality in 191 (97%) of 197 B-cell malignancies and T-cell clonality in 71 (97%) of 73 T-cell malignancies. Limit of detection (LOD) for IG/TCR rearrangements was established at 5% using cell line blends. Chromosomal translocations were detected in 145 (95%) of 152 cases known to be positive. CNAs were validated for immunogenetic and oncogenetic regions, highlighting their novel role in confirming clonality in somatically hypermutated cases. Single-nucleotide variant LOD was determined as 4% allele frequency, and an orthogonal validation using 32 samples resulted in 98% concordance. The EuroClonality-NDC assay is a robust tool providing a single end-to-end workflow for simultaneous detection of B- and T-cell clonality, translocations, CNAs, and sequence variants.
We analyze a time series of global temperature anomaly distributions to identify and estimate persistent features in climate change. We employ a formal test for the existence of functional unit roots in the time series of these densities, and we develop a new test to distinguish functional unit roots from functional deterministic trends or explosive behavior. Results suggest that temperature anomalies contain stochastic trends (as opposed to deterministic trends or explosive roots), two trends are present in the Northern Hemisphere while one stochastic trend is present in the Southern Hemisphere, and the probabilities of observing moderately positive anomalies have increased. We postulate that differences in the pattern and number of unit roots in each hemisphere may be due to a natural experiment which causes human emissions of greenhouse gases and sulfur to be greater in the Northern Hemisphere, decreasing the mean temperature anomaly but increasing the spatial variance relative to the Southern Hemisphere. Together, these results are consistent with the theory of anthropogenic climate change.
Introduction: Current diagnostic standards for lymphoproliferative disorders include detection of clonal immunoglobulin (IG) and/or T cell receptor (TR) rearrangements, translocations, copy number alterations (CNA) and somatic mutations. These analyses frequently require a series of separate tests such as clonality PCR, fluorescence in situ hybridisation and/or immunohistochemistry, MLPA or SNParrays and sequencing. The EuroClonality-NGS DNA capture (EuroClonality-NDC) panel, developed by the EuroClonality-NGS Working Group, was designed to characterise all these alterations by capturing variable, diversity and joining IG and TR genes along with additional clinically relevant genes for CNA and mutation analysis. Methods: Well characterised B and T cell lines (n=14) representing a diverse repertoire of IG/TR rearrangements were used as a proficiency assessment to ensure 7 testing EuroClonality centres achieved optimal sequencing performance using the EuroClonality-NDC optimised and standardised protocol. A set of 56 IG/TR rearrangements across the 14 cell lines were compiled based on detection by Sanger, amplicon-NGS and capture-NGS sequencing technologies. For clinical validation of the NGS panel, clinical samples representing both B and T cell malignancies (n=280), with ≥ 5% tumour infiltration were collected from 10 European laboratories, with 88 (31%) being formalin fixed paraffin-embedded samples. Samples were distributed to the 7 centres for library preparation, hybridisation with the EuroClonality-NDC panel and sequencing on a NextSeq 500, using the EuroClonality-NDC standard protocol. Sequencing data were analysed using a customised version of ARResT/Interrogate, with independent review of the results by 2 centres. All cases exhibiting discordance between the benchmark and capture NGS results were submitted to an internal review committee comprising members of all participating centres. Results: All 7 testing centres detected all 56 rearrangements of the proficiency assessment and continued through to the validation phase. A total of 10/280 (3.5%) samples were removed from the validation analysis due to NGS failures (n=1), tumour infiltration < 5% (n=7), and sample misidentification (n=2). The EuroClonality-NDC panel detected B cell clonality (i.e. detection of at least one clonal rearrangement at IGH, IGK or IGL loci) in 189/197 (96%) B cell malignancies. Seven of the 8 discordant cases were post-germinal centre malignancies exhibiting Ig somatic hypermutation. The EuroClonality-NDC panel detected T cell clonality (i.e. detection of at least one clonal rearrangement at TRA, TRB, TRD or TRG loci) in 70/73 (96%) T cell malignancies. In all 3 discordant cases analysis of benchmark PCR data was not able to detect clonality at any TR loci. Next, we examined whether the EuroClonality-NDC panel could detect clonality at each of the individual loci, resulting in sensitivity values of 95% or higher for all IG/TR loci, with the exception of those where limited benchmark data were available, i.e. IGL (n=3) and TRA (n=7). The specificity of the panel was assessed on benign reactive lesions (n=21) that did not contain clonal IG/TR rearrangements based on BIOMED-2/EuroClonality PCR results; no clonality was observed by EuroClonality-NDC in any of the 21 cases. Limit of detection (LOD) assessment to detect IG/TR rearrangements was performed using cell line blends with each of the 7 centres receiving blended cell lines diluted to 10%, 5.0%, 2.5% and 1.25%. Across all 7 centres the overall detection rate was 100%, 94.1%, 76.5% and 32.4% respectively, giving an overall LOD of 5%. Sufficient data were available in 239 samples for the analysis of translocations. The correct translocation was detected in 137 out of 145 cases, resulting in a sensitivity of 95%. Table 1 shows how translocations identified by the EuroClonality-NDC protocol were restricted to disease subtypes known to harbour those types of translocations. Analysis of CNA and somatic mutations in all samples is underway and will be presented at the meeting. Conclusions: The EuroClonality-NDC panel, with an optimised laboratory protocol and bioinformatics pipeline, detects IG and TR rearrangements and translocations with high sensitivity and specificity with a LOD ≤ 5% and provides a single end-to-end workflow for the simultaneous detection of IG/TR rearrangements, translocations, CNA and sequence variants. Table. Disclosures Stamatopoulos: Janssen: Honoraria, Research Funding; Abbvie: Honoraria, Research Funding. Klapper:Roche, Takeda, Amgen, Regeneron: Honoraria, Research Funding. Ferrero:Gilead: Speakers Bureau; Janssen: Consultancy, Membership on an entity's Board of Directors or advisory committees, Speakers Bureau; EUSA Pharma: Membership on an entity's Board of Directors or advisory committees; Servier: Speakers Bureau. van den Brand:Gilead: Speakers Bureau. Groenen:Gilead: Speakers Bureau. Brüggemann:Incyte: Membership on an entity's Board of Directors or advisory committees; Amgen: Membership on an entity's Board of Directors or advisory committees; Roche: Consultancy. Langerak:Gilead: Research Funding, Speakers Bureau; F. Hoffmann-La Roche Ltd: Research Funding; Genentech, Inc.: Research Funding; Janssen: Speakers Bureau. Gonzalez:Roche: Honoraria, Research Funding; AstraZeneca: Consultancy, Honoraria, Research Funding, Speakers Bureau.
This paper nonparametrically estimates the distribution of world citizens’ income and investigates world income inequality for the period from 1970 to 2010. We consider 188 countries that account for 98.68% of the world population and almost 100% of the world GDP in the year 2010. Various income inequality indices such as the Gini coefficient reveal that the world income inequality dramatically decreased during the 2000s, whereas it only slightly declined from 1970 to 2000, because the across-country inequality component substantially decreased during the 2000s even if the within-country inequality component kept increasing during the 1990s and the 2000s. These findings still hold when we include top income tax data in the analysis. We also propose more sophisticated methods to impute missing top income share data and to combine them with income survey data.
This paper investigates whether the Chinese RMB has become more influential (than the U.S. dollar) in determining the exchange rates of East Asian currencies in recent years. We use a regression method with time-varying coefficients to trace changes in coefficients over time. The empirical results show that the RMB's effects on East Asian currencies were near zero before 2008, but since then have significantly increased and taken over the role of the U.S. dollar in some countries (Indonesia, Malaysia, and the Philippines). In Singapore and Thailand, the RMB is still a non-factor. South Korea shows an interesting pattern, in that the role of the RMB swings over time, with an increase in the past couple of years. We conjecture that the trade share with China has a positive influence on the role of the RMB. In conclusion, given the small absolute value of the regression coefficient on RMB, although the RMB has attained a more significant status in the currency market, it is too early to talk about the creation of an RMB bloc in East Asia.
This paper studies the price elasticity of the peak electricity demand of the residential sector in Korea using high frequency data collected by AMR (Automatic Meter Reading) system. The main purpose of this paper is to estimate the price elasticity by allowing the nonlinear relationship between price and temperature in the short-run residential electricity demand curve. Specifically, we consider a Logistic Smooth Transition Regression model with functional coefficients to capture the temperature-dependent price elasticity of residential peak demand in Korea. We show conclusive evidence that the non-economic variables influence the price elasticity of peak residential demand in Korea. Our estimation results show that the price elasticity is dependent upon temperature, and peak demand becomes more sensitive when the weather is very hot or cold.
This paper proposes a new framework to analyze the nonstationarity in the time series of state densities, representing either cross-sectional or intra-period distributions of some underlying economic variables. We regard each state density as a realization of Hilbertian random variable, and use a functional time series model to fit a given time series of state densities. This allows us to explore various sources of the nonstationarity of such time series. The potential unit roots are identified through functional principal component analysis, and subsequently tested by the generalized eigenvalues of leading components of normalized estimated variance operator. The asymptotic null distribution of the test statistic is obtained and tabulated. We use the methodology developed in the paper to investigate the state densities given by the cross-sectional distributions of individual earnings and the intra-month distributions of stock returns. We find some clear evidence for the presence of strong persistency in their time series.
We propose a novel approach to measure and analyze the short-run effect of temperature on monthly sectoral electricity demand. This effect is specified as a function of the density of temperatures observed at a high frequency with a functional coefficient, in contrast to conventional methods using a function of monthly heating and cooling degree days. Our approach also allows non-climate variables to influence the short-run demand response to temperature changes. Our methodology is demonstrated using Korean electricity demand data for residential and commercial sectors. In the residential sector, we do not find evidence that the non-climate variables affect the demand response to temperature. In contrast, we show conclusive evidence that the non-climate variables influence the demand response in the commercial sector. In particular, commercial consumers are less responsive to cold temperatures when controlling for the electricity price relative to city gas. They are more responsive to the price when temperatures are cold. The estimated effect of the time trend suggests that seasonality of commercial demand has increased in the winter but decreased in the summer.
Published in the Journal of Econometrics (https://doi.org/10.1016/j.jeconom.2019.05.014) We analyze a time series of global temperature anomaly distributions to identify and estimate persistent features in climate change. We employ a formal test for the existence of functional unit roots in the time series of these densities, and we develop a new test to distinguish functional unit roots from functional deterministic trends or explosive behavior. Results suggest that temperature anomalies contain stochastic trends (as opposed to deterministic trends or explosive roots), two trends are present in the Northern Hemisphere while one stochastic trend is present in the Southern Hemisphere, and the probabilities of observing moderately positive anomalies have increased. We postulate that differences in the pattern and number of unit roots in each hemisphere may be due to a natural experiment which causes human emissions of greenhouse gases and sulfur to be greater in the Northern Hemisphere, decreasing the mean temperature anomaly but increasing the spatial variance relative to the Southern Hemisphere. Together, these results are consistent with the theory of anthropogenic climate change. This Version:
BACKGROUND:Gene expression connectivity mapping has gained much popularity recently with a number of successful applications in biomedical research testifying its utility and promise. Previously methodological research in connectivity mapping mainly focused on two of the key components in the framework, namely, the reference gene expression profiles and the connectivity mapping algorithms. The other key component in this framework, the query gene signature, has been left to users to construct without much consensus on how this should be done, albeit it has been an issue most relevant to end users. As a key input to the connectivity mapping process, gene signature is crucially important in returning biologically meaningful and relevant results. This paper intends to formulate a standardized procedure for constructing high quality gene signatures from a user's perspective.RESULTS:We describe a two-stage process for making quality gene signatures using gene expression data as initial inputs. First, a differential gene expression analysis comparing two distinct biological states; only the genes that have passed stringent statistical criteria are considered in the second stage of the process, which involves ranking genes based on statistical as well as biological significance. We introduce a "gene signature progression" method as a standard procedure in connectivity mapping. Starting from the highest ranked gene, we progressively determine the minimum length of the gene signature that allows connections to the reference profiles (drugs) being established with a preset target false discovery rate. We use a lung cancer dataset and a breast cancer dataset as two case studies to demonstrate how this standardized procedure works, and we show that highly relevant and interesting biological connections are returned. Of particular note is gefitinib, identified as among the candidate therapeutics in our lung cancer case study. Our gene signature was based on gene expression data from Taiwan female non-smoker lung cancer patients, while there is evidence from independent studies that gefitinib is highly effective in treating women, non-smoker or former light smoker, advanced non-small cell lung cancer patients of Asian origin.CONCLUSIONS:In summary, we introduced a gene signature progression method into connectivity mapping, which enables a standardized procedure for constructing high quality gene signatures. This progression method is particularly useful when the number of differentially expressed genes identified is large, and when there is a need to prioritize them to be included in the query signature. The results from two case studies demonstrate that the approach we have developed is capable of obtaining pertinent candidate drugs with high precision.
We introduce a panel model with a nonparametric functional coefficient of multiple arguments. The coefficient is a function both of time, allowing temporal changes in an otherwise linear model, and of the regressor itself, allowing nonlinearity. In contrast to a time series model, the effects of the two arguments can be identified using a panel model. We apply the model to the relationship between real GDP and electricity consumption. Our results suggest that the corresponding elasticities have decreased over time in developed countries, but that this decrease cannot be entirely explained by changes in GDP itself or by sectoral shifts.
Quantile normalization (QN) is a technique for microarray data processing and is the default normalization method in the Robust Multi-array Average (RMA) procedure, which was primarily designed for analysing gene expression data from Affymetrix arrays. Given the abundance of Affymetrix microarrays and the popularity of the RMA method, it is crucially important that the normalization procedure is applied appropriately. In this study we carried out simulation experiments and also analysed real microarray data to investigate the suitability of RMA when it is applied to dataset with different groups of biological samples. From our experiments, we showed that RMA with QN does not preserve the biological signal included in each group, but rather it would mix the signals between the groups. We also showed that the Median Polish method in the summarization step of RMA has similar mixing effect. RMA is one of the most widely used methods in microarray data processing and has been applied to a vast volume of data in biomedical research. The problematic behaviour of this method suggests that previous studies employing RMA could have been misadvised or adversely affected. Therefore we think it is crucially important that the research community recognizes the issue and starts to address it. The two core elements of the RMA method, quantile normalization and Median Polish, both have the undesirable effects of mixing biological signals between different sample groups, which can be detrimental to drawing valid biological conclusions and to any subsequent analyses. Based on the evidence presented here and that in the literature, we recommend exercising caution when using RMA as a method of processing microarray gene expression data, particularly in situations where there are likely to be unknown subgroups of samples.
One of the major challenges in systems biology is to understand the complex responses of a biological system to external perturbations or internal signalling depending on its biological conditions. Genome-wide transcriptomic profiling of cellular systems under various chemical perturbations allows the manifestation of certain features of the chemicals through their transcriptomic expression profiles. The insights obtained may help to establish the connections between human diseases, associated genes and therapeutic drugs. The main objective of this study was to systematically analyse cellular gene expression data under various drug treatments to elucidate drug-feature specific transcriptomic signatures. We first extracted drug-related information (drug features) from the collected textual description of DrugBank entries using text-mining techniques. A novel statistical method employing orthogonal least square learning was proposed to obtain drug-feature-specific signatures by integrating gene expression with DrugBank data. To obtain robust signatures from noisy input datasets, a stringent ensemble approach was applied with the combination of three techniques: resampling, leave-one-out cross validation, and aggregation. The validation experiments showed that the proposed method has the capacity of extracting biologically meaningful drug-feature-specific gene expression signatures. It was also shown that most of signature genes are connected with common hub genes by regulatory network analysis. The common hub genes were further shown to be related to general drug metabolism by Gene Ontology analysis. Each set of genes has relatively few interactions with other sets, indicating the modular nature of each signature and its drug-feature-specificity. Based on Gene Ontology analysis, we also found that each set of drug feature (DF)-specific genes were indeed enriched in biological processes related to the drug feature. The results of these experiments demonstrated the potential of the method for predicting certain features of new drugs using their transcriptomic profiles, providing a useful methodological framework and a valuable resource for drug development and characterization.