Elevated urban temperatures are a significant concern across the globe due to their negative health effects and increased energy use. Understanding the spatial variation in urban air temperatures can lead to informed mitigation and planning efforts. Air temperatures for multiple urban areas in the state of Iowa, USA, at three times of the day, were collected using customized sensors mounted on vehicles driven through a variety of landscapes in each urban area. Geographic information systems technology was used to process high-resolution landscape datasets and derive variables that summarize the urban landscape surrounding each temperature measurement point. Five different statistical models: standard regression, trend surface, geostatistical, time series, and random forest, were fitted to nighttime data in the Waterloo–Cedar Falls urban area. We demonstrate that the best method for predicting Waterloo–Cedar Falls nighttime data is to use Waterloo–Cedar Falls data collected at a different time of day. However, when data are not available in the same city for which predicted air temperatures are needed, we explore which substitute city’s data best forecast the target city’s air temperature, via four cross-validation strategies. We find that, when predicting evening and nighttime air temperatures for the Iowa urban areas, choosing the closest-in-population-size substitute city provides the best predicted air temperatures.
A laborshed analysis examines the available workforce that flows from the surrounding communities into a nodal city. Iowa Workforce Development (“IWD”) currently uses employer survey data to create a laborshed study for the largest communities in each of Iowa’s 99 counties. IWD has surveyed 18,428 Iowans from July 2019 to April 2021 to ask how likely these Iowans are to change jobs if they are currently employed, or how apt they are to rejoin the labor force if they are presently unemployed, recently retired or are a homemaker. The likelihood of changing jobs or re-entering the workforce is modeled through a polytomous response logistic regression, using both demographic information and labor market characteristics obtained from the surveyed Iowans. In this study, prediction of individual potential job applicants, together with estimates of workers for each zip code in-commuting to a nodal city, are detailed. In particular, this study estimates the total number of individuals who are eager to change jobs or regain employment, called the Weighted Labor Force (“WLF”), for any desired laborshed in Iowa. The WLF computation is demonstrated for the Cedar Valley Laborshed, which consists of the nodal cities of Waterloo and Cedar Falls in northeastern Iowa.
This case presents a thorough exposition of the state of the art in the calculation and interpretation of automated valuation model (AVM) performance metrics, including the forecast standard deviation (FSD), confidence scores, vertical and horizontal equity, and error buckets. It also discusses the failure rate, failure magnitude, and failure median absolute percentage error (MAPE) metrics, which focus on tails of an AVM's distribution of errors. In addition, this case demonstrates the calculations and relationships between the AVM performance metrics using a regression model and property sales from a medium-sized Midwestern college city.
This work proposes a hedonic random field model to describe house selling prices from 2000 to 2005 in Cedar Falls, Iowa. This real estate market presents two distinctive features that are not well described by traditional stationary Gaussian random field models: (a) the city has, on its periphery, a hoglot that acts as an externality, affecting both the mean and variance of the selling prices, and (b) the distribution of house selling prices display heavy tails, even after the distance to the hoglot and house-specific covariates are accounted for in the mean structure of the model. A non-stationary and non-Gaussian random field model is constructed by multiplying two independent Gaussian random fields tailored to model the probabilistic features displayed by the Cedar Falls dataset. A Markov chain Monte Carlo algorithm that uses data augmentation is employed to fit the proposed model.
Some households are willing to pay a premium to live farther away from a disamenity, such as a neighborhood gas station with a leaking underground storage tank (LUST). In this study, submarkets are constructed to allow for this premium to be a function of both the distance to the nearest LUST and the intensity of multiple LUSTs. We find evidence that households within ¼ mile of multiple LUSTs do not have a statistically significant aversion to living near them. However, households more than a ¼ mile from a LUST are willing to pay 9.29% more for a house located 10% farther away. (Q51, Q53, R21)
In most studies, standardized test scores are used as a proxy for school quality. Standardized test scores, however, may not fully capture the value of a public school to the households who live in the school’s attendance zone. We use the sudden closure of a well-performing public school in Iowa to estimate this value. Holding other things constant, we find that the school added 6.8% (about $9,000 for the mean house price) to the value of houses in the attendance zone over and above any effect associated with standardized test scores.
A point source, non-stationary covariance structure model is proposed, having only one additional parameter over a standard, stationary covariance structure, spatial model. Additionally, the proposed model is demonstrated to fit better than the three extra parameter, point source, non-stationary spatial model proposed by Ecker and De Oliveira (Commun Stat Theory Methods 37:2066–2078, 2008 ). The proposed model is fit from a Bayesian perspective and illustrated using a house sales dataset from Cedar Falls, Iowa.
This work proposes a non stationary random field model to describe the spatial variability of housing prices that are affected by a localized externality. The model allows for the effect of the localized externality on house prices to be represented in the mean function and/or the covariance function of the random field. The correlation function of the proposed model is a mixture of an isotropic correlation function and a correlation function that depends on the distances between home sales and the localized externality. The model is fit using a Bayesian approach via a Markov chain Monte Carlo algorithm. A dataset of 437 single family home sales during 2001 in the city of Cedar Falls, Iowa, is used to illustrate the model.
The impact of 39 swine confined or concentrated animal feeding operations (CAFOs) in Black Hawk County, Iowa on 5,822 house sales is explored by introducing a new variable that more accurately captures the effects of prevailing winds, exploring potential adverse effects within concentric circles around each CAFO, managing selection bias, and incorporating spatial correlation into the error term of the empirical model. Large adverse impacts suffered by houses that are within 3 miles and directly downwind from a CAFO are found. Beyond 3 miles, CAFOs have a generally decreasing adverse impact on house prices as distance to the CAFO increases.
This study develops and fits a nonlinear, unified convex–concave model for land sales. The model's flexibility accommodates convexity for small parcels, concavity for large parcels, a non-deterministic change point while accounting for spatial correlation. All parameters are fit simultaneously from a Bayesian perspective using Markov Chain Monte Carlo techniques. Virtually all previous models for land sales are shown to be special cases of the unified model. The results indicate an 86.9% probability of convexity (plottage) for parcels smaller than 3516 ft2.
This project examines two Iowa lakes to explore the feasibility of using remote sensing technologies for assessing water quality in lieu of actual ground samples. We demonstrate that a principal component analysis of the more than 20,000 remote sensed pixels can be used in a regression analysis to accurately predict total phosphorus levels in Casey Lake on three distinct times in the summer of 2004.
It is well established that house prices are dynamic. It is also axiomatic that location influences such selling prices, motivating our objective of incorporating spatial information in explaining the evolution of house prices over time. In this paper, we propose a rich class of spatio-temporal models under which each property is point referenced and its associated selling price modeled through a collection of temporally indexed spatial processes. Such modeling includes and extends all house price index models currently in the literature, and furthermore permits distinction between the effects of time and location. We study single family residential sales in two distinct submarkets of a metropolitan area and further categorize the data into single- and multiple-transaction observations. We find the spatial component is very important in explaining house price. Moreover, the relative homogeneity of homes within the submarket and the frequency with which homes sell affects the pattern of variation across space and time. Differences between single and repeat sale data are evident. The methodology is applicable to more general capital asset pricing when location is anticipated to be influential.
We examine the evolution of the National Basketball Association’s (NBA’s) Draft Lottery by showing how the changes implemented by the Board of Governors have impacted the probabilities of obtaining top picks in the ensuing draft lottery. We explore how these changes have impacted the team with the worst record and also investigate the conditional probabilities of the fourth worst team receiving the third pick. These calculations are conditioned upon two specified teams receiving the first two selections. We show that the probability of the fourth worst team receiving the third pick can be made unconditionally. We calculate the probabilities for the fourth worst team to move up in the draft to receive either the first, second, or third selections, along with its chance of keeping the fourth pick or even dropping in the draft. We find there is a higher chance for the fourth worst team to drop to the fifth, sixth or seventh position than to stay at the fourth position or move up.
For modeling spatial processes, we propose a rich parametric class of stationary range anisotropic covariance structures that, when applied in R 2 , greatly increases the scope of variogram contors. Geometric anisotropy, which provides the most common generalization of isotropy within stationarity, is a special case. Our class is built from monotonic isotropic correlation functions and special cases include the Matérn and the general exponential functions. As a result, our range anisotropic correlation specification can be attached to a second order stationary spatial process model, unlike ad hoc approaches to range anisotropy in the literature. We adopt a Bayesian perspective to obtain full inference and demonstrate how to fit the resulting model using sampling-based methods. In the presence of measurement error/microscale effect, we can obtain both the usual predictive as well as the noiseless predictive distribution. We analyze a data set of scallop catches under the general exponential range anisotropic model, withholding ten sites to compare the accuracy and precision of the standard and noiseless predictive distributions.
Water quality has become and important issue in the state of Iowa as well as across the entire United States. Two Iowa Lakes, Silver Lake and Casey Lake were chosen for study by a team of biologists, chemists, earth scientists and statisticians from the University of Northern Iowa. Our goals are to statistically compare the water quality in the two lakes in each year and examine whether or not each lake has changed, in terms of water quality variables, from 1999 to 2000. In addition, we explore which variables most affect phosphorus levels in each lake in 2000. Lastly, we explore the spatial distribution of phosphorus in the sediment of each lake. Discriminant Analyses and ANCOVA show significant difference between the two lakes in both 1999 and 2000 as well as a change in Silver Lake's water quality data from 1999 to 2000. Regression Analyses show that, in Silver Lake, phosphorus levels increased during the summer of 2000 while they decreased with increasing levels of surface dissolved oxygen and decreased as the water became less clear. The analyses also show that phosphorus levels in Lake Casey decreased as the water became less clear. A significant relationship between phosphorus in the sediment and depth exists in Lake Casey. While a significant 2-dimensional spatial correlation cannot be shown in Silver Lake, spatial analyses do show the existence of a significant 3-dimensional spatial correlation in Lake Casey.
Extreme concentrations of water quality variables can cause serious adverse effects in an ecosystem, making their detection an important environmental issue. In Chesapeake Bay, a decreasing gradient of total nitrogen concentration extends from the highest values in the north at the mouth of the Susquehanna river to the lowest values in the south near the Atlantic ocean. We propose a general definition of ‘hot spot’ that includes previous definitions and is appealing for processes with a spatial trend. We model these data using the Bayesian Transformed Gaussian (BTG) random field model proposed by De Oliveira et al . ( 1997 ), which combines the Box–Cox family of power transformations and a spatial trend. The median function is used as the measure of spatial trend, which offers some advantages over the customarily used mean function. The BTG model is fitted by an enhanced Monte Carlo algorithm, and the methodology is applied to the nitrogen concentration data. Copyright © 2002 John Wiley & Sons, Ltd.
This paper blends and extends the Colwell and Munneke (1997) and Isakson (1997) models of urban land values to include a spatiotemporal plattage component. By including a monocentric term in the Isakson model, we arrive at a Colwell and Munneke-type model with a temporal component included in the plattage term. We further extend this model to include a spatiotemporal effect in the plattage term and demonstrate the use of geostatistical plattage models. (R14)
The problem of modeling and analyzing point referencedbinary spatial data is addressed. We formulate a hierarchical model introducingspatial eects at the second stage. Rather than capturing thesecond stage spatial association using a Markov random eld specication,we employ a second order stationary Gaussian spatial process. In fact,we introduce a convenient latent Gaussian spatial process making our observeddata a realization of an indicator process on this latent process.Convenient...
A geometrically anisotropic spatial process can be viewed as being a linear transformation of an isotropic spatial process. Customary semivariogram estimation techniques often involve ad hoc selection of the linear transformation to reduce the region to isotropy and then fitting a valid parametric semivariogram to the data under the transformed coordinates. We propose a Bayesian methodology which simultaneously estimates the linear transformation and the other semivariogram parameters. In addition, the Bayesian paradigm allows full inference for any characteristic of the geometrically anisotropic model rather than merely providing a point estimate. Our work is motivated by a dataset of scallop catches in the Atlantic Ocean in 1990 and also in 1993. The 1990 data provide useful prior information about the nature of the anisotropy of the process. Exploratory data analysis (EDA) techniques such as directional empirical semivariograms and the rose diagram are widely used by practitioners. We recommend a suitable contour plot to detect departures from isotropy. We then present a fully Bayesian analysis of the 1993 scallop data, demonstrating the range of inferential possibilities.
For modeling spatial processes, we introduce a rich class of range anisotropic covariance structures which greatly increases the scope of variogram contours and includes geometric anisotropy and isotropy as special cases. Spatial aspects can also be captured using geographical covariates to create a trend surface. We adopt a Bayesian perspective to study models involving both trend surface and range anisotropic covariance parameters, tting these models using sampling-based methods. We analyze a data set of scallop catches and withhold ten sites to compare the accuracy and precision of a noiseless version of the predictive distribution for two such parametric models. We also estimate the detrended variogram and, for a particular subregion of interest, predict locally and globally on the original scale of the data. .