The frequency of the first digits of numbers drawn from an exponential probability density oscillate around the Benford frequencies. Analysis, simulations and empirical evidence show that datasets must have at least 10,000 entries for these oscillations to emerge from finite-sample noise. Anecdotal evidence from population data is provided.
Algorithms for identifying public health threats or disease outbreaks are vulnerable to false alarms arising from sudden shifts in health‐care utilization or data participation. This paper describes a method of reducing false alerts in automated public health surveillance algorithms, and in particular, automated syndromic surveillance algorithms, that rely on health‐care utilization data. The technique is based on monitoring syndromic counts with reference to a suitable background, or reference, series of counts. The suitability of the background time series in decreasing the false‐alarm rate will be shown to be related mathematically to the so‐called mutual information that exists between the random variables representing the syndromic and background time series of counts. The method can be understood as a noise cancellation filter technique in which one noisy (reference) channel is used to cancel the background noise of the monitored (measured) channel. The issues discussed here may also be relevant to the appropriate use of rates in epidemiology and biostatistics. Copyright © 2011 John Wiley & Sons, Ltd.
utomated systems for public health surveillance have evolved over the past several years as national and local institutions have been learning the most effective ways to share and apply population health information for outbreak investigation and tracking. The changes have included developments in algorithmic alerting methodology. This article presents research efforts at The Johns Hopkins University Applied Physics Laboratory for advancing this methodology. The analytic methods presented cover outcome variable selection, background estimation, determination of anomalies for alerting, and practical evaluation of detection performance. The methods and measures are adapted from information theory, signal processing, financial forecasting, and radar engineering for effective use in the biosurveillance data environment. Examples are restricted to univariate algorithms for daily time series of syndromic data, with discussion of future generalization and enhancement.
An approach to identifying public health threats by characterizing syndromic surveillance data in terms of its surprisability is discussed. Surprisability in our model is measured by assigning a probability distribution to a time series, and then calculating its entropy, leading to a straightforward designation of an alert. Initial application of our method is to investigate the applicability of using suitably-normalized syndromic counts (i.e., proportions) to improve early event detection.
Use of a high-order deterministic sampling technique in direct simulation Monte-Carlo (DSMC) simulations eliminates statistical noise and improves computational performance by orders of magnitude. In this paper it is also shown that if a random timestep is used in place of a fixed timestep, there is an additional improvement in performance. This performance can be increased by using a timestep that samples a random variable with a high-kurtosis probability density function. As a simple example of the method, the one-dimensional diffusion equation with an exponentially-distributed timestep is simulated, and a performance gain of approximately two is obtained. Applications to numerical simulations of fluids and plasmas are indicated.
[1] We want to thank Rietveld and Stubbe [2006] (hereinafter referred to as RS) for providing us with the opportunity to clarify the issues addressed by our work, including the differences in experimental results and theoretical interpretation with the early pioneering work of Rietveld et al. [1986, 1987, 1989] and Rietveld and Stubbe [1987]. We first apologize for the inadvertent omission of Rietveld et al. [1986] that occurred accidentally when the accepted manuscript was shortened to fit the GRL length guidelines. The particular figure we refer to in our paper is Figure 1 of Rietveld et al. [1986] (hereinafter referred to as R86), that was repeated as Figure 2 of Rietveld and Stubbe [1987]. With this we want to address the more substantive issues: [2] 1. We agree with RS that when disagreements between theory and experiment arise the theory should be questioned. For this reason we conducted our short pulse experiment. Figure 1a shows the magnetic response of the electrojet to 2.5 msec heating reproduced from Figure 1 of R86 along with our own experimental result (Figure 1b). The traces are similar with one critical exception: the approximately 1 pT field that follows turn on and extends from approximately 0.6 to 3 msecs in Figure 1b. RS admit that this feature was absent from their measurements and attribute it to the processing they used to convert their loop antenna measurements to a time plot for the magnetic field. This explains why Figure 1a does not represent the correct magnetic response of the electrojet to short pulse heating. Note that our experimental result shown in Figure 1b is consistent with the theoretical prediction derived using a three-dimensional Green's function. [3] 2. In our paper we referred to Figure 1b as three-component with asymmetry for pulse on vs. off. This is schematically shown in Figure 2 that approximates the response as an impulse with duration To, found to be of the order of .25 msec for the experimental conditions, superimposed on a square pulse with duration equal to the pulse length and amplitude a fraction of the maximum amplitude of the pulse (in this case approximately .25). During the off time there is only an impulse with reverse polarity but no magnetic plateau component, resulting in asymmetry. This is different than the well-known heating/cooling asymmetry explored in R86. The field structure shown in Figure 2 and revealed for the first time during the HAARP campaign is crucial in resolving the efficiency vs. frequency puzzle documented in Figure 8 of Rietveld et al. [1989], as well as revealing a possible flaw in their treatment of the issue. [4] 3. Rietveld and Stubbe's previous work does not seem to have resolved many critical issues concerning the scaling of the efficiency with frequency. Consider their results shown in Figure 3 (corresponding to Figure 8 of Rietveld et al. [1989]). As seen, the theoretical curve (solid line) is in complete disagreement with the data. [5] Furthermore the more than ten dB lower efficiency in the range below 1 kHz cannot be accounted for theoretically. A major point of our paper was to explain the puzzling efficiency variations shown in Figure 3. Note that the amplitude peaks at about 2 kHz, due to effects attributable to the dependence of the observed waveforms on frequency or equivalently pulse length. The frequency scaling for frequencies below 2 kHz (corresponding to To > .25 msec) can be found by taking the Fourier transform of the waveform shown in Figure 2 that approximates the observed waveforms. The results are shown in Figure 4 for two values of the ratio α of the maximum amplitude of the square pulse to that of the impulse. It reproduces well both the sharp reduction below 2.0 kHz and the expected flattening at lower frequencies. For frequencies higher than 2 kHz, and neglecting the presence of minor enhancements at frequencies 2, 4, 6 and 8 kHz (that were first correctly attributed by Rietveld et al. [1989] to transit time in-phase resonance with the ionospheric reflection) the decrease in the amplitude can be attributed to the fact that for times T < To the modified conductivity has not yet reached saturation as shown in Figure 1 of Papadopoulos et al. [2005]. [6] 4. The final point addresses the differences in our respective modeling approaches and their critique of our theory. In general both models follow a similar path to the point of calculating the spatiotemporal profile of the currents induced by the conductivity modifications. From then on, the two approaches diverge. As noted by Rietveld et al. [1989, p. 273] they used the steady state electron temperature enhancement to compute the modified conductivity. Following that, they used full wave antenna theory assuming plane wave propagation to compute the magnetic field amplitude as a function of frequency. As explained in Papadopoulos et al. [2005], we used the spatiotemporal profile of the current induced by the conductivity modification in a three dimensional Green's function to find the waveform on the ground. The model used by Rietveld et al. [1989] to compute the theoretical scaling shown in Figure 3 does not incorporate the physics involved with the time derivative of the conductivity for two reasons: [7] a. It uses the steady state value of the conductivity. [8] b. The plane wave approximation is equivalent to a one-dimensional approximation. As noted in Morse and Feshbach [1953, pp. 867–868] the one dimensional Green's function of equation (1) of our paper results in a magnetic field amplitude simply proportional to the current. It is only in a three dimensional analysis (the appropriate one for the problem under consideration) that a magnetic field amplitude proportional to the time derivative of the current is introduced. [9] We hope that the above remarks clarify the issues raised by Rietveld and Stubbe in their comment and establish the importance of this work in resolving critical issues of ELF/VLF generation efficiency. [10] This work was supported by the Office of Naval Research contract N00014-03-C-0466.