In this article, we explore patterns of R-2 across data sets to reveal a fundamental equivalence. If any set of bivariate data is rotated, the R-2 value changes and defines a curve as a function of the rotation angle theta. We show that for any data set whose rotated R-2 curve has the same maximum value, the entire curve is identical up to a shift in the rotation angle. We also find that a recently introduced measure of linearity, Q(2), provides a definitive relation between these data sets, with equal Q(2) values resulting in identically rotated R-2 curves. Finally, we also show that we can consider R-2 as the combination of two components, the linearity of the data and a rotation angle. While rotating data or axes for non spatial data loses the specific meaning of the variables, these results may provide additional understanding of the relationships within the data and can reveal similarities between data sets, the role of sensitivity in correlation, and the nature of correlation itself.
Abstract The simple linear regression model and the associated goodness-of-fit measure, the coefficient of determination, R 2, are only appropriate when all measurement errors are associated with the measurement of the data in the dependent variable. When measurement errors are assumed in both variables, a Deming regression can be used; however, there is no associated R 2-type measure for this specific type of regression. In this paper, we propose a measure, which utilizes the minimum percentage improvement of the Deming regression over either the horizontal or the vertical line through the centroid of the data. We investigate some properties of this measure and its relation to R 2. We also consider other candidate methodologies for a generalized R 2 measure for a Deming regression model and investigate strengths and weaknesses of each as a way of beginning the conversation about which measure is the best and for which applications it is most suited.
Humans possess a remarkable ability to recognize both simple patterns such as shapes and handwriting and very complex patterns such as faces and landscapes. To investigate one small aspect of human pattern recognition, in this study participants position lines of "best fit" to two-dimensional scatter plots of data. The study investigates the variation in participants' fits and whether there is some consistent metric being used in fitting the lines. For example, is there a natural tendency toward fitting lines similarly to one of the standard regression lines: vertical, horizontal, or orthogonal. This study also investigates the effect of outliers on the line a participant fits to a scatter plot with a strong linear trend and provides guidance for future inquiries.
Rounding is a necessary step in many mathematical processes. We are taught early in our education about significant figures and how to properly round a number. So when we are given a data set and asked to find a regression line, we are inclined to offer the line with rounded coefficients to reflect our model. However, the effects are not as insignificant as they might seem at first. In this paper, we investigate some consequences of rounding the coefficients in a least squares linear regression with respect to the calculated value of R2, and consider ways to minimize the amount of error that can arise.
You have lost contact with an unmanned surveillance plane as it is flying over a large stretch of uninhabited desert. You send high-altitude reconnaissance aircraft to take pictures of a potential ...