We propose PACE-GGM, a data-adaptive differentially private method for covariance estimation that concentrates its privacy budget on the most informative entries of the empirical covariance matrix, rather than perturbing all entries. This applies in the natural setting where the modeler supplies separate bounds for each variable, so that individual entries can be measured with less noise than the full matrix. In each round, our method selects a poorly approximated entry, measures it using the Gaussian mechanism, and then reconstructs a full covariance matrix using a maximum-entropy reconstruction objective, leading to a Gaussian graphical model structure. Experiments on diverse real-world datasets demonstrate consistent improvements in estimation error with respect to the Gaussian mechanism and other baselines, particularly in high-dimensional and low-to-moderate privacy regimes.
Background Accurate information on population-level movements of migratory animals is essential for understanding migration and for designing effective conservation strategies in a changing world. Yet such information remains scarce for most migratory species due to the effort and expense needed to collect data across their full distribution ranges. BirdFlow is a probabilistic modeling framework that infers population-level movements from weekly species distribution maps produced by the participatory science project eBird. However, BirdFlow models have only been tuned for a handful of species using high-resolution individual tracking data, which is not available for most migratory species. Methods Here, we introduce a general tuning and evaluation framework for BirdFlow that enables the first large-scale integration of distributional and individual-level data to infer animal movement across continents and hundreds of migratory species, eliminating reliance on any single individual-tracking data source. By generalizing the BirdFlow model parametrization, we enable tuning and validation using multiple complementary data sources, including GPS tracks, banding recoveries, and radio telemetry data from the Motus Wildlife Tracking System. We investigate the efficacy of this approach by (1) investigating predictive performance compared to null models; (2) validating the biological plausibility of BirdFlow models by comparing movement properties such as route straightness, number of stopovers, and migration speed between model-generated routes and real movement tracks; and (3) comparing the performance of models tuned on species-specific movement data to models tuned using hyperparameters transferred from other species. Results Our results show that BirdFlow models produced by the new tuning framework achieve biologically realistic performance, even for prediction horizons of thousands of kilometers and several months. When species-specific data are unavailable, models can still be tuned using data from other phylogenetically adjacent species to achieve improved performance. Conclusions By integrating eBird Status & Trends abundance surfaces with data from banding recaptures, radio telemetry, and GPS tracking, we scale BirdFlow models to 153 North American migratory species, representing the first collection of continental-scale population-level movement and forecasting models. Species-specific tuning improves population-level movement forecasts, while taxonomically informed hyperparameter transfer supports the modeling of data-limited species. Overall, our work offers a foundation for more accurate predictions across hundreds of species for research in ecology and conservation, disease surveillance, aviation, and public outreach.
Computer vision models are increasingly used as measurement tools to estimate population-level quantities from large image collections, but prediction errors introduce bias and the resulting estimates lack statistical guarantees required in scientific applications. Prior work uses a Monte Carlo framework to combine model predictions with ground-truth annotations by sampling some images for humans to label and is able to provide unbiased estimates with controllable accuracy, but primarily addresses single-scalar estimation. We study the more general problem of multi-target estimation, where many quantities (e.g., class counts or proportions) must be estimated simultaneously, and adapt sampling and estimation strategies from survey sampling to this setting. Evaluations on five detection and segmentation datasets with 7-80 classes show that importance sampling excels with moderate annotation budgets or fewer targets, whereas uniform sampling with control variates is superior when estimating many targets or operating with minimal labels. Additionally, a subset-based ratio estimator remains highly competitive across all regimes. Ultimately, our framework effectively combines biased model predictions and limited human labels into rigorous scientific measurements.
Two-point correlation functions (2PCF) are widely used to characterize how points cluster in space. In this work, we study the problem of measuring the 2PCF over a large set of points, restricted to a subset satisfying a property of interest. An example comes from astronomy, where scientists measure the 2PCF of star clusters, which make up only a tiny subset of possible sources within a galaxy. This task typically requires careful labeling of sources to construct catalogs, which is time-consuming. We present a human-in-the-loop framework for efficient estimation of 2PCF of target sources. By leveraging a pre-trained classifier to guide sampling, our approach adaptively selects the most informative points for human annotation. After each annotation, it produces unbiased estimates of pair counts across multiple distance bins simultaneously. Compared to simple Monte Carlo approaches, our method achieves substantially lower variance while significantly reducing annotation effort. We introduce a novel unbiased estimator, sampling strategy, and confidence interval construction that together enable scalable and statistically grounded measurement of two-point correlations in astronomy datasets.
Privately releasing marginals of a tabular dataset is a foundational problem in differential privacy. However, state-of-the-art mechanisms suffer from a computational bottleneck when marginal estimates are reconstructed from noisy measurements. Recently, residual queries were introduced and shown to lead to highly efficient reconstruction in the batch query answering setting. We introduce new techniques to integrate residual queries into state-of-the-art adaptive mechanisms such as AIM. Our contributions include a novel conceptual framework for residual queries using multi-dimensional arrays, lazy updating strategies, and adaptive optimization of the per-round privacy budget allocation. Together these contributions reduce error, improve speed, and simplify residual query operations. We integrate these innovations into a new mechanism (AIM+GReM), which improves AIM by using fast residual-based reconstruction instead of a graphical model approach. Our mechanism is orders of magnitude faster than the original framework and demonstrates competitive error and greatly improved scalability.
Aim: Quantifying movement patterns of migratory birds throughout their annual cycles is critical for effective conservation planning. The BirdFlow project uses eBird participatory science data products to model species-level migratory trajectories across entire ranges. However, BirdFlow trajectory models are not readily compatible with location-based monitoring technologies such as weather radar. We develop a scalable, species-level measurement of nocturnal bird migration traffic and assess its utility for integration with radar monitoring. Innovation: We introduce the BirdFlow migration traffic rate (BMTR), a location-based metric of migratory passage for individual species derived from BirdFlow models. BMTR quantifies the weekly proportion of a species' population passing over a transect, can be calculated across a species' entire range and can be converted to absolute numbers using population estimates. We applied this framework to 312 nocturnal migratory bird species in North America. Aggregating all species' migration activities, we compared them with average weekly nocturnal migration traffic detected by 152 NEXRAD weather surveillance radars over 28 years. We further evaluated the use of BMTR to disaggregate radar-derived migration traffic into species-level estimates. Main Conclusions: BirdFlow- and radar-derived migration traffic showed strong agreement (r = 0.784), particularly along the Mississippi and Atlantic Flyways, with models incorporating demographic adjustments performing best. Annual nocturnal migration traffic estimated by BirdFlow and radar was closely aligned, with BirdFlow estimates averaging 33.3% higher. BMTR effectively captured major flyways, high-traffic routes and seasonal migration dynamics. It also enabled disaggregation of radar traffic into species-level contributions, revealing dominant species at specific locations. BMTR provides a new methodological bridge between trajectory-based BirdFlow movement models and location-based monitoring approaches. It offers a scalable, species-specific tool for migration ecology and conservation, supports monitoring in areas without radar coverage and enables new opportunities for public engagement through visualisation of near real-time, species-level migration.
Long‐term monitoring of bird populations across scales is important in evaluating conservation targets and creating effective conservation strategies. For nearly six decades, the Breeding Bird Survey (BBS) has served as the primary broad‐scaled source of relative abundance trends of swallows and martins in North America. Recently, however, it has become possible to obtain breeding population trends using semi‐structured eBird community science data. Moreover, weather surveillance radar data of swallow and martin roosting populations yield a third complementary source of trend information. Using results from these three approaches, we propose a novel method of spatially combining estimates of percent change per year into a probability of directional agreement and/or disagreement that describes (1) the direction of the trend within a given region, (2) the amount of evidence associated with the estimate and (3) how much uncertainty surrounds it. We focus our efforts on an area of high Hirundinidae concentration in the North American Great Lakes region and predict trends from 2012 to 2022. We found a high probability of agreement between all three sources about observed declines in swallow and martin trends in the region surrounding Lake Ontario and to the west of Lake Michigan. Focusing future research on these regions could improve our understanding of these declines and help build more targeted conservation initiatives. Synthesis and applications . Our data integration methodology allows managers to identify regions that accumulate evidence of concerning trends across multiple wildlife monitoring schemes. These regions can thus be prioritized in conservation and management efforts. This approach can be generalized to other sources of long‐term monitoring data of different species, at different stages of their annual cycle, in any geographic location.
Abstract The US NEXRAD radar network has monitored the aerosphere over the US and its territories continuously since the 1990s and archived nearly 300 million radar volume scans. These data contain a wealth of information about the movements of birds, bats, and insects. Historically, this biological information was difficult to access due to the amount of data and challenges in analyzing it. In the last 15 years, fueled by computational and methodological advances, large-scale aeroecology research has blossomed. However, comprehensive analyses of the NEXRAD archive remain very costly. We collected measurements from every volume scan in the NEXRAD archive—nearly 300 million data files total—to assemble a dataset of aerial biological activity over the US from 1995 to 2025. The core data are vertical profiles, which summarize biological activity at different heights above the radar station for each volume scan. We also provide time series data products that aggregate vertical profiles to point measurements at radar stations across time. These data products can support a range of aeroecology analyses at significantly reduced effort.
Sufficient statistic perturbation (SSP) is a widely used method for differentially private linear regression. SSP adopts a data-independent approach where privacy noise from a simple distribution is added to sufficient statistics. However, sufficient statistics can often be expressed as linear queries and better approximated by data-dependent mechanisms. In this paper we introduce data-dependent SSP for linear regression based on post-processing privately released marginals, and find that it outperforms state-of-the-art data-independent SSP. We extend this result to logistic regression by developing an approximate objective that can be expressed in terms of sufficient statistics, resulting in a novel and highly competitive SSP approach for logistic regression. We also make a connection to synthetic data for machine learning: for models with sufficient statistics, training on synthetic data corresponds to data-dependent SSP, with the overall utility determined by how well the mechanism answers these linear queries.
Aim: Migratory birds are under threat by climate change. Successfully conserving them requires knowing which climatic factors drive changes in their migratory behaviour. Weather conditions may directly or indirectly affect the temporally disjointed life history stages of migratory birds, including the breeding, roosting and nonbreeding stages. However, the influences of these broad-scale patterns are often not studied together. Coupling migratory bird movements estimated using weather radar (NEXRAD) with long-term and large-scale environmental data allows us to overcome these spatiotemporal uncertainties. Here, we assess environmental drivers of the phenology of postbreeding roosting of aerial insectivores in the Great Lakes region (USA) by evaluating predictors during the months leading up to roosting across species' ranges. Location: Northern United States and Canada. Time Period: 21-year (2000-2020). Major Taxa Studied: Avian aerial insectivores. Methods: We conducted a spatially explicit time-window analysis to examine the effects of 17 gridded weather and vegetation variables on swallow peak roosting phenology in the Great Lakes, making minimal ecological assumptions. ResultsWe found that peak roosting timing is paced by both local conditions (headwind at 850 hPa) at the Great Lakes and distant conditions (minimum temperature, precipitation rate and specific humidity) at the likely breeding and stopover sites, with warmer temperatures advancing, headwind delaying and high precipitation advancing the phenophases. Time windows selected for the possible breeding and stopover sites are mostly before or around the time of roosting, with one exception during wintertime. Main Conclusions: Although climatic shifts play a significant role in driving variation in phenology, for migratory species, the proximate driver can originate hundreds to thousands of kilometres away, and potentially months prior. Our study illuminates these far-reaching patterns in aerial insectivores, enhancing our grasp of migration ecology and paving the way for a more comprehensive understanding of hemispheric animal movements.
Earth's lower atmosphere is a vital ecological habitat, home to trillions of organisms that live, forage, and migrate through this medium. Despite its importance, this space is seldom considered a primary habitat for ecological or conservation prioritization, making it one of the least studied environments. However, it plays a crucial role as a global conduit for the transfer of biomass, weather, and inorganic materials. Fundamental research is essential to address core ecological questions related to the ecological consequences of this habitat's intricate spatial and temporal structure. To advance our understanding of airspace use by migratory animals, we analyzed over 108 million 5-min radar observations from 143 NEXRAD sites, focusing on 24-h diel cycles across the contiguous United States. This extensive dataset, spanning from 1995 to 2022, allowed us to quantify aerial space use by systematically identifying peak activity times, the portion of the airspace that contained the majority of migration activity, and the percentage of migrants passing across diurnal and nocturnal diel cycles. We found that airspace is used predominantly during nocturnal periods in both spring and autumn (88%), while summer exhibited a more balanced distribution (54% nocturnal). Additionally, the percentage of nocturnal activity increased with latitude in spring and autumn but decreased in summer. Peak aerial activity typically occurred about 4 h after local sunset in both spring and autumn, with variations based on latitude and longitude. During these peak times, on average, half of the aerial movement was confined within a vertical band of 516 meters, starting around 355 m above ground level. Our research underscores the need to view the lower atmosphere as a structured habitat with significant ecological importance.
Migrating landbirds adjust their flight and stopover behaviors to efficiently cross inhospitable geographies, such as the Gulf of Mexico and the Sahara Desert. In addition to these natural barriers, birds may increasingly encounter anthropogenic barriers created by large-scale changes in land use. One such barrier could be the Corn Belt in the Midwest United States, where 76.4% of precolonial vegetation (forest and grassland combined) has been replaced by agricultural and urban areas, primarily corn fields. We used 5 years of data from 47 weather radar stations in the United States to compare the population-level flight patterns of migrating landbirds crossing the Corn Belt and the forested landscapes south and north of it in spring and autumn. We also examined the impacts of the Corn Belt relative to the Gulf of Mexico on the stopover behavior of migrating birds by comparing changes in the proportion of migrants that stop to rest (stopover-to-passage ratio [SPR]) relative to distance from both barriers. Birds showed increased meridional airspeeds and stronger selection for tailwinds when crossing the Corn Belt compared with forested landscapes. For birds crossing the Gulf of Mexico, the highest proportion of migrants stopped to rest after crossing the Gulf, and SPR decreased sharply as distance from the shoreline increased. We did not find this pattern after migrants crossed the Corn Belt, although the SPR increased in the Corn Belt as birds approached the down-route forest boundary in both seasons. This weaker pattern for stopover propensity after crossing the Corn Belt is likely due to its narrower width, the availability of small forest patches throughout the Corn Belt, and the subset of species affected, compared with the gulf. We recommend restoring stepping stones of forest in the Corn Belt and protecting woodlands along the Gulf Coast to help landbirds successfully negotiate both barriers.
The widespread availability of off-the-shelf machine learning models poses a challenge: which model, of the many available candidates, should be chosen for a given data analysis task? This question of model selection is traditionally answered by collecting and annotating a validation dataset – a costly and time-intensive process. We propose a method for active model selection, using predictions from candidate models to prioritize the labeling of test data points that efficiently differentiate the best candidate. Our method, CODA, performs consensus-driven active model selection by modeling relationships between classifiers, categories, and data points within a probabilistic framework. The framework uses the consensus and disagreement between models in the candidate pool to guide the label acquisition process, and Bayesian inference to update beliefs about which model is best as more information is collected. We validate our approach by curating a collection of 26 benchmark tasks capturing a range of model selection scenarios. CODA outperforms existing methods for active model selection significantly, reducing the annotation effort required to discover the best model by upwards of 70
AI has the potential to transform scientific discovery by analyzing vast datasets with little human effort. However, current workflows often do not provide the accuracy or statistical guarantees that are needed. We introduce active measurement, a human-in-the-loop AI framework for scientific measurement. An AI model is used to predict measurements for individual units, which are then sampled for human labeling using importance sampling. With each new set of human labels, the AI model is improved and an unbiased Monte Carlo estimate of the total measurement is refined. Active measurement can provide precise estimates even with an imperfect AI model, and requires little human effort when the AI model is very accurate. We derive novel estimators, weighting schemes, and confidence intervals, and show that active measurement reduces estimation error compared to alternatives in several measurement tasks.
During their nonbreeding period, many species of swallows and martins (family: Hirundinidae) congregate in large communal roosts. Some of these roosts are well-known within local birdwatching communities; however, monitoring them at large spatial scales and with day-to-day temporal resolution is challenging. Community-science platforms such as the Purple Martin Conservation Association's project MartinRoost and eBird have addressed some of these challenges by centralizing data collected from regional communities. Additionally, due to the high densities of birds within these aggregations, their early morning dispersals are systematically detected by weather radars, which have also been used to collect data about roost timing and location. An important issue, however, limits spatiotemporal scope of previous radar-based studies: finding the roost signatures on millions of rendered reflectivity images is extremely time-consuming. Recent advances in computer vision, however, have allowed us to reduce this effort. The rise of this technology makes it necessary that we assess whether our biological definition of a roost matches what the machine-learning models are capturing. We do so by comparing eBird detections of roosts in the Great Lakes region with those obtained by a human-supervised machine-learning model from 2000 to 2022. With more than two decades of data, we assess the ability of these two tools to detect roosts on a day-to-day basis, and we compare the phenology of dispersals to investigate whether radar detections correspond to swallow and martin roosts or if they are associated with other well-known birds that form large aggregations. Our comparison of these datasets strongly suggests that swallows and martins are responsible for the dispersals we observe on the radars from July to late September; however, the alternative species we examined could be causing some of the detections in October.
The ecological pressures that maintain the behavioral preferences of avian migrants, such as the timing and duration of nocturnal flights, remain elusive yet are critical to understand the evolution of the migratory program. In this study, we use an atypical light condition – extremely short to non‐existent nights at high latitudes – to study responses to a forced tradeoff between nocturnality and migratory flight duration. Through this lens, we aim to elucidate the relative importance of the pressures shaping the dynamics of migratory flights. We use next‐generation radar (NEXRAD) data from seven stations across Alaska to characterize the timing of peak migratory activity relative to dusk, the duration of elevated migratory activity, and the fraction of migratory activity falling within the night. We find that as night lengths decrease, the timing of peak migration clusters tightly around solar midnight, resulting in peak activity occurring slightly closer to dusk. Meanwhile, the duration of elevated migratory activity, while becoming less variable, is largely maintained, resulting in a major shift towards diurnal migration as night lengths decrease below the average duration of nightly migratory activity. These results demonstrate that the stabilizing pressure on the times during which migrants fly is strong and generally overrides the pressure of maintaining nocturnal migration, contextualizing geographic variation in migratory dynamics across Alaska, and elucidating the structure of decisions that determine migratory behavior.
Computer vision-based re-identification (Re-ID) systems are increasingly being deployed for estimating population size in large image collections. However, the estimated size can be significantly inaccurate when the task is challenging or when deployed on data from new distributions. We propose a human-in-the-loop approach for estimating population size driven by a pairwise similarity derived from an off-the-shelf Re-ID system. Our approach, based on nested importance sampling, selects pairs of images for human vetting driven by the pairwise similarity, and produces asymptotically unbiased population size estimates with associated confidence intervals. We perform experiments on various animal Re-ID datasets and demonstrate that our method outperforms strong baselines and active clustering approaches. In many cases, we are able to reduce the error rates of the estimated size from around 80\% using CV alone to less than 20\% by vetting a fraction (often less than 0.002\%) of the total pairs. The cost of vetting reduces with the increase in accuracy and provides a practical approach for population size estimation within a desired tolerance when deploying Re-ID systems.
Many applications use computer vision to detect and count objects in massive image collections. However, automated methods may fail to deliver accurate counts, especially when the task is very difficult or requires a fast response time. For example, during disaster response, aid organizations aim to quickly count damaged buildings in satellite images to plan relief missions, but pre-trained building and damage detectors often perform poorly due to domain shifts. In such cases, there is a need for human-in-the-loop approaches to accurately count with minimal human effort. We propose DISCount -- a detector-based importance sampling framework for counting in large image collections. DISCount uses an imperfect detector and human screening to estimate low-variance unbiased counts. We propose techniques for counting over multiple spatial or temporal regions using a small amount of screening and estimate confidence intervals. This enables end-users to stop screening when estimates are sufficiently accurate, which is often the goal in real-world applications. We demonstrate our method with two applications: counting birds in radar imagery to understand responses to climate change, and counting damaged buildings in satellite imagery for damage assessment in regions struck by a natural disaster. On the technical side we develop variance reduction techniques based on control variates and prove the (conditional) unbiasedness of the estimators. DISCount leads to a 9-12x reduction in the labeling costs to obtain the same error rates compared to naive screening for tasks we consider, and surpasses alternative covariate-based screening approaches.