Diffusion models represent a new paradigm in text-to-image generation. Beyond generating high-quality images from text prompts, models such as Stable Diffusion have been successfully extended to the joint generation of semantic segmentation pseudo-masks. However, current extensions primarily rely on extracting attentions linked to prompt words used for image synthesis. This approach limits the generation of segmentation masks derived from word tokens not contained in the text prompt. In this work, we introduce Open-Vocabulary Attention Maps (OVAM)-a training-free method for text-to-image diffusion models that enables the generation of attention maps for any word. In addition, we propose a lightweight optimization process based on OVAM for finding tokens that generate accurate attention maps for an object class with a single annotation. We evaluate these tokens within existing state-of-the-art Stable Diffusion extensions. The best-performing model improves its mIoU from 52.1 to 86.6 for the synthetic images' pseudo-masks, demonstrating that our optimized tokens are an efficient way to improve the performance of existing methods without architectural changes or retraining.
We previously reported that single doses of the norepinephrine transporter inhibitor, atomoxetine, increased standing blood pressure (BP) and ameliorated symptoms in patients with neurogenic orthostatic hypotension (nOH). We aimed to evaluate the effect of atomoxetine over four weeks in patients with nOH. A randomized, double-blind, placebo-controlled crossover clinical trial between July 2016 and May 2021 was carried out with an initial open-label, single-dose phase (10 or 18 mg atomoxetine), followed by a 1-week wash-out, and a subsequent double-blind 4-week treatment sequence (period 1: atomoxetine followed by placebo) or vice versa (period 2). The trial included a 2-week wash-out period. The primary endpoint was symptoms of nOH as measured by the orthostatic hypotension questionnaire (OHQ) assessed at 2 weeks. A total of 68 patients were screened, 40 were randomized, and 37 completed the study. We found no differences in the OHQ composite score between atomoxetine and placebo at 2 weeks (−0.3 ± 1.7 versus −0.4 ± 1.5; P = 0.806) and 4 weeks (−0.6 ± 2.4 versus −0.5 ± 1.6; P = 0.251). There were no differences either in the OHSA scores at 2 weeks (3 ± 1.9 versus 4 ± 2.1; P = 0.062) and at 4 weeks (3 ± 2.2 versus 3 ± 2.0; P = 1.000) or in the OH daily activity scores (OHDAS) at 2 weeks (4 ± 3.0 versus 5 ± 3.1, P = 0.102) and 4 weeks (4 ± 3.0 versus 4 ± 2.7, P = 0.095). Atomoxetine was well-tolerated. While previous evidence suggested that acute doses of atomoxetine might be efficacious in treating nOH; results of this clinical trial indicated that it was not superior to placebo to ameliorate symptoms of nOH. ClinicalTrials.gov; NCT02316821.
This paper introduces UrbAM-ReID, a new long-term geo-positioned urban ReID dataset. It is composed by four subdatasets recording the same trajectory at the UAM Campus, each one recorded in different seasons and including an inverse direction recording. While most of the current datasets in the state-of-the-art focus on person re-identification, with vehicles as the second most explored object, our work specifically addresses urban objects re-identification, currently, waste containers, rubbish bins, and crosswalks. The dataset provides different attributes of the annotated objects, like their classes, their foreground or background status and the geo-position. Several evaluation configurations can be defined to simulate realistic scenarios that may arise in actual situations within the management of urban elements, considering the utilization of just visual data, or incorporating additional attributes, providing different complexity levels. Finally, the dataset is used for defining a benchmark where two state-of-the-art systems are evaluated. The dataset and supplementary material is available in https://github.com/vpulab/UrbAMReID
Quantifying space use and segregation, as well as the extrinsic and intrinsic factors affecting them, is crucial to increase our knowledge of species-specific movement ecology and to design effective management and conservation measures. This is particularly relevant in the case of species that are highly mobile and dependent on sparse and unpredictable trophic resources, such as vultures. Here, we used the GPS-tagged data of 127 adult Griffon Vultures Gyps fulvus captured at five different breeding regions in Spain to describe the movement patterns (home-range size and fidelity, and monthly cumulative distance). We also examined how individual sex, season, and breeding region determined the cumulative distance traveled and the size and overlap between consecutive monthly home-ranges. Overall, Griffon Vultures exhibited very large annual home-range sizes of 5027 +/- 2123 km(2), mean monthly cumulative distances of 1776 +/- 1497 km, and showed a monthly home-range fidelity of 67.8 +/- 25.5%. However, individuals from northern breeding regions showed smaller home-ranges and traveled shorter monthly distances than those from southern ones. In all cases, home-ranges were larger in spring and summer than in winter and autumn, which could be related to difference in flying conditions and food requirements associated with reproduction. Moreover, females showed larger home-ranges and less monthly fidelity than males, indicating that the latter tended to use the similar areas throughout the year. Overall, our results indicate that both extrinsic and intrinsic factors modulate the home-range of the Griffon Vulture and that spatial segregation depends on sex and season at the individual level, without relevant differences between breeding regions in individual site fidelity. These results have important implications for conservation, such as identifying key threat factors necessary to improve management actions and policy decisions.
There is a knowledge gap in the study of Argasidae soft ticks and the pathogens they can transmit. These hematophagous arthropods are widely distributed and are often considered typical bird ectoparasites. Tick-parasitized birds can act not only as a reservoir of pathogens but also can carry these pathogen-infected arthropods to new areas. Seven griffon vulture nestlings were sampled in northeastern Spain, collecting ticks ( n = 28) from two individuals and blood from each vulture ( n = 7). Blood samples from griffon vultures tested PCR positive for Flavivirus (7/7), Anaplasma (6/7), piroplasms (4/7), and Rickettsia (1/7). A total of 27 of the 28 analyzed ticks were positive for Rickettsia , 9/28 for Anaplasma , 2/28 for piroplasms, and 5/28 for Crimean-Congo Hemorrhagic fever virus (CCHFv). Sequencing and phylogenetic analyses confirmed the presence of Rickettsia spp., Babesia ardeae , and zoonotic Anaplasma phagocytophilum in vultures and Rickettsia spp., B. ardeae , and CCHFv genotype V in ticks.
The main goal of vehicle re-identification (ReID) is to associate the same vehicle identity in different cameras. This is a challenging task due to variations in light, viewpoints or occlusions; in particular, vehicles present a large intra-class variability and a small inter-class variability. In ReID, the samples in the test sets belong to identities that have not been seen during training. To reduce the domain gap between train and test sets, this work explores unsupervised domain adaptation (UDA) generating automatically pseudo-labels from the testing data, which are used to fine-tune the ReID models. Specifically, the pseudo-labels are obtained by clustering using different hyperparameters and incrementally due to retraining the model a number of times per hyperparameter with the generated pseudo-labels. The ReID system is evaluated in CityFlow-ReID-v2 dataset.
In semantic segmentation, training data down-sampling is commonly performed due to limited resources, the need to adapt image size to the model input, or improve data augmentation. This down-sampling typically employs different strategies for the image data and the annotated labels. Such discrepancy leads to mismatches between the down-sampled color and label images. Hence, the training performance significantly decreases as the down-sampling factor increases. In this paper, we bring together the down-sampling strategies for the image data and the training labels. To that aim, we propose a novel framework for label down-sampling via soft-labeling that better conserves label information after down-sampling. Therefore, fully aligning soft-labels with image data to keep the distribution of the sampled pixels. This proposal also produces reliable annotations for under-represented semantic classes. Altogether, it allows training competitive models at lower resolutions. Experiments show that the proposal outperforms other down-sampling strategies. Moreover, state-of-the-art performance is achieved for reference benchmarks, but employing significantly less computational resources than foremost approaches. This proposal enables competitive research for semantic segmentation under resource constraints.
Vehicle re-identification (ReID) aims to find a specific vehicle identity across multiple non-overlapping cameras. The main challenge of this task is the large intra-class and small inter-class variability of vehicles appearance, sometimes related with large viewpoint variations, illumination changes or different camera resolutions. To tackle these problems, we proposed a vehicle ReID system based on ensembling deep learning features and adding different post-processing techniques. In this paper, we improve that proposal by: incorporating large-scale synthetic datasets in the training step; performing an exhaustive ablation study showing and analyzing the influence of synthetic content in ReID datasets, in particular CityFlow-ReID and VeRi-776; and extending post-processing by including different approaches to the use of gallery video-clips of the target vehicles in the re-ranking step. Additionally, we present an evaluation framework in order to evaluate CityFlow-ReID: as this dataset has not public ground truth annotations, AI City Challenge provided an on-line evaluation service which is no more available; our evaluation framework allows researchers to keep on evaluating the performance of their systems in the CityFlow-ReID dataset.
Cross-camera image data association is essential for many multi-camera computer vision tasks, such as multi-camera pedestrian detection, multi-camera multi-target tracking, 3D pose estimation, etc. This association task is typically modeled as a bipartite graph matching problem and often solved by applying minimum-cost flow techniques, which may be computationally demanding for large data. Furthermore, cameras are usually treated by pairs, obtaining local solutions, rather than finding a global solution at once for all multiple cameras. Other key issue is that of the affinity function: the widespread usage of non-learnable pre-defined distances, such as the Euclidean and Cosine ones. This paper proposes an effective approach for cross-camera data-association focused on a global solution, instead of processing cameras by pairs. To avoid the usage of fixed distances and thresholds, we leverage the connectivity of Graph Neural Networks, previously unused in this scope, using a Message Passing Network to jointly learn features and similarity functions. We validate the proposal for pedestrian cross-camera association, showing results over the EPFL multi-camera pedestrian dataset. Our approach considerably outperforms the literature data association techniques, without requiring to be trained in the same scenario in which it is tested. Our code is available at https://www-vpu.eps.uam.es/publications/gnn
The widespread use of second-generation anticoagulant rodenticides (SGARs) and their high persistence in an-imal tissues has led to these compounds becoming ubiquitous in rodent-predator-scavenger food webs. Exposure to SGARs has usually been investigated in wildlife species found dead, and despite growing evidence of the potential risk of secondary poisoning of predators and scavengers, the current worldwide exposure of free-living scavenging birds to SGARs remains scarcely investigated. We present the first active monitoring of blood SGAR concentrations and prevalence in the four European obligate (i.e., vultures) and facultative (red and black kites) avian scavengers in NE Spain. We analysed 261 free-living birds and detected SGARs in 39.1% (n = 102) of individuals. Both SGAR prevalence and concentrations (ESGARs) were related to the age and foraging behaviour of the species studied. Black kites showed the highest prevalence (100%), followed by red kites (66.7%), Egyptian (64.2%), bearded (20.9%), griffon (16.9%) and cinereous (6.3%) vultures. Overall, both the prevalence and average ESGARs were higher in non-nestlings than nestlings, and in species such as kites and Egyptian vultures foraging in anthropic landscapes (e.g., landfill sites and livestock farms) and exploiting small/medium-sized carrions. Brodifacoum was most prevalent (28.8%), followed by difenacoum (16.1%), flocoumafen (12.3%) and bromadiolone (7.3%). In SGAR-positive birds, the ESGAR (mean & PLUSMN; SE) was 7.52 & PLUSMN; 0.95 ng mL-1; the highest level detected being 53.50 ng mL-1. The most abundant diastereomer forms were trans-bromadiolone and flo-coumafen, and cis-brodifacoum and difenacoum, showing that lower impact formulations could reduce sec-ondary exposures of non-target species. Our findings suggest that SGARs can bioaccumulate in scavenging birds, showing the potential risk to avian scavenging guilds in Europe and elsewhere. We highlight the need for further studies on the potential adverse effects associated with concentrations of SGARSs in the blood to better interpret active monitoring studies of free-living birds.
Background Multiple system atrophy (MSA) is a fatal neurodegenerative disease characterized by the aggregation of alpha-synuclein in glia and neurons. Sirolimus (rapamycin) is an mTOR inhibitor that promotes alpha-synuclein autophagy and reduces its associated neurotoxicity in preclinical models. Objective To investigate the efficacy and safety of sirolimus in patients with MSA using a futility design. We also analyzed 1-year biomarker trajectories in the trial participants. Methods Randomized, double-blind, parallel group, placebo-controlled clinical trial at the New York University of patients with probable MSA randomly assigned (3:1) to sirolimus (2-6 mg daily) for 48 weeks or placebo. Primary endpoint was change in the Unified MSA Rating Scale (UMSARS) total score from baseline to 48 weeks. ( NCT03589976). Results The trial was stopped after a pre-planned interim analysis met futility criteria. Between August 15, 2018 and November 15, 2020, 54 participants were screened, and 47 enrolled and randomly assigned (35 sirolimus, 12 placebo). Of those randomized, 34 were included in the intention-to-treat analysis. There was no difference in change from baseline to week 48 between the sirolimus and placebo in UMSARS total score (mean difference, 2.66; 95% CI, -7.35-6.91; P = 0.648). There was no difference in UMSARS-1 and UMSARS-2 scores either. UMSARS scores changes were similar to those reported in natural history studies. Neuroimaging and blood biomarker results were similar in the sirolimus and placebo groups. Adverse events were more frequent with sirolimus. Analysis of 1-year biomarker trajectories in all participants showed that increases in blood neurofilament light chain (NfL) and reductions in whole brain volume correlated best with UMSARS progression. Conclusions Sirolimus for 48 weeks was futile to slow the progression of MSA and had no effect on biomarkers compared to placebo. One-year change in blood NfL and whole brain atrophy are promising biomarkers of disease progression for future clinical trials. (c) 2022 International Parkinson and Movement Disorder Society
Despite its increasing recognition and extensive research, there is no unifying hypothesis on the pathophysiology of the postural tachycardia syndrome. In this cross-sectional study, we examined the role of fear conditioning and its association with tachycardia and cerebral hypoperfusion on standing in 28 patients with postural tachycardia syndrome (31 ± 12 years old, 25 females) and 21 matched controls. We found that patients had higher somatic vigilance (P = 0.0167) and more anxiety (P < 0.0001). They also had a more pronounced anticipatory tachycardia right before assuming the upright position in a tilt-table test (P = 0.015), a physiological indicator of fear conditioning to orthostasis. While standing, patients had faster heart rate (P < 0.001), higher plasma catecholamine levels (P = 0.020), lower end-tidal CO2 (P = 0.005) and reduced middle cerebral artery blood flow velocity (P = 0.002). Multi-linear logistic regression modelling showed that both epinephrine secretion and excessive somatic vigilance predicted the magnitude of the tachycardia and the hyperventilation. These findings suggest that the postural tachycardia syndrome is a functional disorder in which standing may acquire a frightful quality, so that even when experienced alone it may elicit a fearful conditioned response. Heightened somatic anxiety is associated with and may predispose to a fear-conditioned hyperadrenergic state when standing. Our results have therapeutic implications.
Multi-Target Multi-Camera (MTMC) vehicle tracking is an essential task of visual traffic monitoring, one of the main research fields of Intelligent Transportation Systems. Several offline approaches have been proposed to address this task; however, they are not compatible with real-world applications due to their high latency and post-processing requirements. In this paper, we present a new low-latency online approach for MTMC tracking in scenarios with partially overlapping fields of view (FOVs), such as road intersections. Firstly, the proposed approach detects vehicles at each camera. Then, the detections are merged between cameras by applying cross-camera clustering based on appearance and location. Lastly, the clusters containing different detections of the same vehicle are temporally associated to compute the tracks on a frame-by-frame basis. The experiments show promising low-latency results while addressing real-world challenges such as the a priori unknown and time-varying number of targets and the continuous state estimation of them without performing any post-processing of the trajectories.
This letter focuses on the task of Multi-Target Multi-Camera vehicle tracking. We propose to associate single-camera trajectories into multi-camera global trajectories by training a Graph Convolutional Network. Our approach simultaneously processes all cameras providing a global solution, and it is also robust to large cameras unsynchronizations. Furthermore, we design a new loss function to deal with class imbalance. Our proposal outperforms the related work showing better generalization and without requiring ad-hoc manual annotations or thresholds, unlike compared approaches.
This paper describes the scientific achievements of a collaboration between a research group and the waste management division of a company. While these results might be the basis for several practical or commercial developments, we here focus on a novel scientific contribution: a methodology to automatically generate geo-located waste container maps. It is based on the use of Computer Vision algorithms to detect waste containers and identify their geographic location and dimensions. Algorithms analyze a video sequence and provide an automatic discrimination between images with and without containers. More precisely, two state-of-the-art object detectors based on deep learning techniques have been selected for testing, according to their performance and to their adaptability to an on-board real-time environment: EfficientDet and YOLOv5. Experimental results indicate that the proposed visual model for waste container detection is able to effectively operate with consistent performance disregarding the container type (organic waste, plastic, glass and paper recycling,…) and the city layout, which has been assessed by evaluating it on eleven different Spanish cities that vary in terms of size, climate, urban layout and containers’ appearance.
OBJECTIVES:To build a model of local hospital utilization resulting from SARS-CoV-2 and to continuously update it with new data.STUDY DESIGN:Retrospective analysis of real performance resulting from a model deployed in a major regional health system.METHODS:Using hospitalization data from the Kaiser Permanente Mid-Atlantic States integrated care system during the period from March 10, 2020, through December 31, 2020, and a custom-developed genetic particle filtering algorithm, we modeled the SARS-CoV-2 outbreak in the mid-Atlantic region. This model produced weekly forecasts of COVID-19-related hospital admissions, which we then compared with actual hospital admissions over the same period.RESULTS:We found that the model was able to accurately capture the data-generating process (weekly mean absolute percentage error, 10.0%-48.8%; Anderson-Darling P value of .97 when comparing percentiles of observed admissions with the uniform distribution) once the effects of social distancing could be accurately measured in mid-April. We also found that our estimates of key parameters, including the reproductive rate, were consistent with consensus literature estimates.CONCLUSIONS:The genetic particle filtering algorithm that we have proposed is effective at modeling hospitalizations due to SARS-CoV-2. The methods used by our model can be reproduced by any major health care system for the purposes of resource planning, staffing, and population care management to create an effective forecasting regimen at scale.
OBJECTIVE:To construct and publicly release a set of medical concept embeddings for codes following the ICD-10 coding standard which explicitly incorporate hierarchical information from medical codes into the embedding formulation. MATERIALS AND METHODS:We trained concept embeddings using several new extensions to the Word2Vec algorithm using a dataset of approximately 600,000 patients from a major integrated healthcare organization in the Mid-Atlantic US. Our concept embeddings included additional entities to account for the medical categories assigned to codes by the Clinical Classification Software Revised (CCSR) dataset. We compare these results to sets of publicly released pretrained embeddings and alternative training methodologies. RESULTS:We found that Word2Vec models which included hierarchical data outperformed ordinary Word2Vec alternatives on tasks which compared naïve clusters to canonical ones provided by CCSR. Our Skip-Gram model with both codes and categories achieved 61.4% normalized mutual information with canonical labels in comparison to 57.5% with traditional Skip-Gram. In models operating on two different outcomes, we found that including hierarchical embedding data improved classification performance 96.2% of the time. When controlling for all other variables, we found that co-training embeddings improved classification performance 66.7% of the time. We found that all models outperformed our competitive benchmarks. DISCUSSION:We found significant evidence that our proposed algorithms can express the hierarchical structure of medical codes more fully than ordinary Word2Vec models, and that this improvement carries forward into classification tasks. As part of this publication, we have released several sets of pretrained medical concept embeddings using the ICD-10 standard which significantly outperform other well-known pretrained vectors on our tested outcomes.
Vehicle re-identification has the objective of finding a specific vehicle among different vehicle crops captured by multiple cameras placed at multiple intersections. Among the different difficulties, high intra-class variability and high inter-class similarity can be highlighted. Moreover, the resolution of the images can be different, which also means a challenge in the re-identification task. Intending to face these problems, we use as baseline our previous work based on obtaining different deep learning features and ensembling them to get a single, stable and robust feature vector. It also includes post-processing techniques that explode all the information provided by the CityFlowV2-ReID dataset, including a re-ranking step. Then, in this paper, several newly included improvements are described. Background and orientation similarity matrices are added to the system to reduce bias towards these characteristics. Furthermore, we take into account the camera labels to penalize the gallery images that share camera with the query image. Additionally, to improve the training step, a synthetic dataset is added to the original one.
Objective Attention networks learn an intelligent weighted averaging mechanism over a series of entities, providing increases to both performance and interpretability. In this article, we propose a novel time-aware transformer-based network and compare it to another leading model with similar characteristics. We also decompose model performance along several critical axes and examine which features contribute most to our model's performance. Materials and methods Using data sets representing patient records obtained between 2017 and 2019 by the Kaiser Permanente Mid-Atlantic States medical system, we construct four attentional models with varying levels of complexity on two targets (patient mortality and hospitalization). We examine how incorporating transfer learning and demographic features contribute to model success. We also test the performance of a model proposed in recent medical modeling literature. We compare these models with out-of-sample data using the area under the receiver-operator characteristic (AUROC) curve and average precision as measures of performance. We also analyze the attentional weights assigned by these models to patient diagnoses. Results We found that our model significantly outperformed the alternative on a mortality prediction task (91.96% AUROC against 73.82% AUROC). Our model also outperformed on the hospitalization task, although the models were significantly more competitive in that space (82.41% AUROC against 80.33% AUROC). Furthermore, we found that demographic features and transfer learning features which are frequently omitted from new models proposed in the EMR modeling space contributed significantly to the success of our model. Discussion We proposed an original construction of deep learning electronic medical record models which achieved very strong performance. We found that our unique model construction outperformed on several tasks in comparison to a leading literature alternative, even when input data was held constant between them. We obtained further improvements by incorporating several methods that are frequently overlooked in new model proposals, suggesting that it will be useful to explore these options further in the future.