Computer vision is transforming fashion industry through Virtual Try-On (VTON) and Virtual Try-Off (VTOFF). VTON generates images of a person in a specified garment using a target photo and a standardized garment image, while a more challenging variant, Person-to-Person Virtual Try-On (p2p-VTON), uses a photo of another person wearing the garment. VTOFF, in contrast, extracts standardized garment images from photos of clothed individuals. We introduce Multi-Garment TryOffDiff (MGT), a diffusion-based VTOFF model capable of handling diverse garment types, including upper-body, lower-body, and dresses. MGT builds on a latent diffusion architecture with SigLIP-based image conditioning to capture garment characteristics such as shape, texture, and pattern. To address garment diversity, MGT incorporates class-specific embeddings, achieving state-of-the-art VTOFF results on VITON-HD and competitive performance on DressCode. When paired with VTON models, it further enhances p2p-VTON by reducing unwanted attribute transfer, such as skin tone, ensuring preservation of person-specific characteristics. Demo, code, and models are available at: https://rizavelioglu.github.io/tryoffdiff/
In the life cycle of highly automated systems operating in an open and dynamic environment, the ability to adjust to emerging challenges is crucial. For systems integrating data-driven AI-based components, rapid responses to deployment issues require fast access to related data for testing and reconfiguration. In the context of automated driving, this especially applies to road obstacles not included in the training data, commonly referred to as out-of-distribution (OoD) road obstacles. Given the availability of large uncurated driving scene recordings, a pragmatic approach is to query a database to retrieve similar scenarios featuring the same safety concerns due to OoD road obstacles. In this work, we extend beyond identifying OoD road obstacles in video streams and offer a comprehensive approach to extract sequences of OoD road obstacles using text queries, thereby proposing a way of curating a collection of OoD data for subsequent analysis. Our proposed method leverages the recent advances in OoD segmentation and multi-modal foundation models to identify and efficiently extract safety-relevant scenes from unlabeled videos. We present a first approach for the novel task of text-based OoD object retrieval, which addresses the question "Have we ever encountered this before?".
We prove several universal approximation results at minimal or near-minimal width for approximation of L^p(ℝ^d_x, ℝ^d_y) and C^0(ℝ^d_x, ℝ^d_y) on compact sets. Our approach uses a unified coding scheme that yields explicit constructions relying only on standard analytic tools. We show that feedforward neural networks with two leaky ReLU activations σ_α, σ_-α achieve the optimal width max{d_x, d_y} for L^p approximation, while a single leaky ReLU σ_α achieves width max{2, d_x, d_y}, providing an alternative proof of the results of Cai et al. (2023). By generalizing to stepped leaky ReLU activations, we extend these results to uniform approximation of continuous functions while identifying sets of activation functions compatible with gradient-based training. Since our constructions pass through an intermediate dimension of one, they imply that autoencoders with a one-dimensional feature space are universal approximators. We further show that squashable activations combined with FLOOR achieve width max{3, d_x, d_y} for uniform approximation. We also establish a lower bound of max{d_x, d_y} + 1 for networks when all activations are continuous and monotone and d_y ≤ 2d_x. Moreover, we extend our results to invertible LU-decomposable networks, proving distributional universal approximation for LU-Net normalizing flows and providing a constructive proof of the classical theorem of Brenier and Gangbo on L^p approximation by diffeomorphisms.
In the realm of fashion object detection and segmentation for online shopping images, existing state-of-the-art fashion parsing models encounter limitations, particularly when exposed to non-model-worn apparel and close-up shots. To address these failures, we introduce FashionFail; a new fashion dataset with e-commerce images for object detection and segmentation. The dataset is efficiently curated using our novel annotation tool that leverages recent foundation models. The primary objective of FashionFail is to serve as a test bed for evaluating the robustness of models. Our analysis reveals the shortcomings of leading models, such as Attribute-Mask R-CNN and Fashionformer. Additionally, we propose a baseline approach using naive data augmentation to mitigate common failure cases and improve model robustness. Through this work, we aim to inspire and support further research in fashion item detection and segmentation for industrial applications. The dataset, annotation tool, code, and models are available at https://rizavelioglu.github.io/fashionfail/.
This paper introduces Virtual Try-Off (VTOFF), a novel task generating standardized garment images from single photos of clothed individuals. Unlike Virtual Try-On (VTON), which digitally dresses models, VTOFF extracts canonical garment images, demanding precise reconstruction of shape, texture, and complex patterns, enabling robust evaluation of generative model fidelity. We propose TryOffDiff, adapting Stable Diffusion with SigLIP-based visual conditioning to deliver high-fidelity reconstructions. Experiments on VITON-HD and Dress Code datasets show that TryOffDiff outperforms adapted pose transfer and VTON baselines. We observe that traditional metrics such as SSIM inadequately reflect reconstruction quality, prompting our use of DISTS for reliable assessment. Our findings highlight VTOFF's potential to improve e-commerce product imagery, advance generative model evaluation, and guide future research on high-fidelity reconstruction. Demo, code, and models are available at: https://rizavelioglu.github.io/tryoffdiff
Deep neural networks (DNN) have made impressive progress in the interpretation of image data so that it is conceivable and to some degree realistic to use them in safety critical applications like automated driving. From an ethical standpoint, the AI algorithm should take into account the vulnerability of objects or subjects on the street that ranges from “not at all”, e.g. the road itself, to “high vulnerability” of pedestrians. One way to take this into account is to define the cost of confusion of one semantic category with another and use cost-based decision rules for the interpretation of probabilities, which are the output of DNNs. However, it is an open problem how to define the cost structure, who should be in charge to do that, and thereby define what AI-algorithms will actually “see”. As one possible answer, we follow a participatory approach and set up an online survey to ask the public to define the cost structure. We present the survey design and the data acquired along with an evaluation that also distinguishes between perspective (car passenger vs. external traffic participant) and gender. Using simulation based F-tests, we find highly significant differences between the groups. These differences have consequences on the reliable detection of pedestrians in a safety critical distance to the self-driving car. We discuss the ethical problems that are related to this approach and also discuss the problems emerging from human–machine interaction through the survey from a psychological point of view. Finally, we include comments from industry leaders in the field of AI safety on the applicability of survey based elements in the design of AI functionalities in automated driving.
We present a human-in-the-loop dashboard tailored to diagnosing potential spurious features that NLI models rely on for predictions. The dashboard enables users to generate diverse and challenging examples by drawing inspiration from GPT-3 suggestions. Additionally, users can receive feedback from a trained NLI model on how challenging the newly created example is and make refinements based on the feedback. Through our investigation, we discover several categories of spurious correlations that impact the reasoning of NLI models, which we group into three categories: Semantic Relevance, Logical Fallacies, and Bias. Based on our findings, we identify and describe various research opportunities, including diversifying training data and assessing NLI models' robustness by creating adversarial test suites.
In this work we present two video test data sets for the novel computer vision (CV) task of out of distribution tracking (OOD tracking). Here, OOD objects are understood as objects with a semantic class outside the semantic space of an underlying image segmentation algorithm, or an instance within the semantic space which however looks decisively different from the instances contained in the training data. OOD objects occurring on video sequences should be detected on single frames as early as possible and tracked over their time of appearance as long as possible. During the time of appearance, they should be segmented as precisely as possible. We present the SOS data set containing 20 video sequences of street scenes and more than 1000 labeled frames with up to two OOD objects. We furthermore publish the synthetic CARLA-WildLife data set that consists of 26 video sequences containing up to four OOD objects on a single frame. We propose metrics to measure the success of OOD tracking and develop a baseline algorithm that efficiently tracks the OOD objects. As an application that benefits from OOD tracking, we retrieve OOD sequences from unlabeled videos of street scenes containing OOD objects.
LU-Net is a simple and fast architecture for invertible neural networks (INN) that is based on the factorization of quadratic weight matrices A=LU, where L is a lower triangular matrix with ones on the diagonal and U an upper triangular matrix. Instead of learning a fully occupied matrix A, we learn L and U separately. If combined with an invertible activation function, such layers can easily be inverted whenever the diagonal entries of U are different from zero. Also, the computation of the determinant of the Jacobian matrix of such layers is cheap. Consequently, the LU architecture allows for cheap computation of the likelihood via the change of variables formula and can be trained according to the maximum likelihood principle. In our numerical experiments, we test the LU-Net architecture as generative model on several academic datasets. We also provide a detailed comparison with conventional invertible neural networks in terms of performance, training as well as run time.
Semantic segmentation is a crucial component for perception in automated driving. Deep neural networks (DNNs) are commonly used for this task and they are usually trained on a closed set of object classes appearing in a closed operational domain. However, this is in contrast to the open world assumption in automated driving that DNNs are deployed to. Therefore, DNNs necessarily face data that they have never encountered previously, also known as anomalies, which are extremely safety-critical to properly cope with. In this work, we first give an overview about anomalies from an information-theoretic perspective. Next, we review research in detecting semantically unknown objects in semantic segmentation. We demonstrate that training for high entropy responses on anomalous objects outperforms other recent methods, which is in line with our theoretical findings. Moreover, we examine a method to assess the occurrence frequency of anomalies in order to select anomaly types to include into a model's set of semantic categories. We demonstrate that these anomalies can then be learned in an unsupervised fashion, which is particularly suitable in online applications based on deep learning.
Bringing deep neural networks (DNNs) into safety critical applications such as automated driving, medical imaging and finance, requires a thorough treatment of the model's uncertainties. Training deep neural networks is already resource demanding and so is also their uncertainty quantification. In this overview article, we survey methods that we developed to teach DNNs to be uncertain when they encounter new object classes. Additionally, we present training methods to learn from only a few labels with help of uncertainty quantification. Note that this is typically paid with a massive overhead in computation of an order of magnitude and more compared to ordinary network training. Finally, we survey our work on neural architecture search which is also an order of magnitude more resource demanding then ordinary network training.
Motor neuron disorders (MND) include a group of pathologies that affect upper and/or lower motor neurons. Among them, amyotrophic lateral sclerosis (ALS) is characterized by progressive muscle weakness, with fatal outcomes only in a few years after diagnosis. On the other hand, primary lateral sclerosis (PLS), a more benign form of MND that only affects upper motor neurons, results in life-long progressive motor dysfunction. Although the outcomes are quite different, ALS and PLS present with similar symptoms at disease onset, to the degree that both disorders could be considered part of a continuum. These similarities and the lack of reliable biomarkers often result in delays in accurate diagnosis and/or treatment. In the nervous system, lipids exert a wide variety of functions, including roles in cell structure, synaptic transmission, and multiple metabolic processes. Thus, the study of the absolute and relative concentrations of a subset of lipids in human pathology can shed light into these cellular processes and unravel alterations in one or more pathways. In here, we report the lipid composition of longitudinal plasma samples from ALS and PLS patients initially, and after 2 years following enrollment in a clinical study. Our analysis revealed common aspects of these pathologies suggesting that, from the lipidomics point of view, PLS and ALS behave as part of a continuum of motor neuron disorders.
Using an integrative, multi-tissue design, we sought to characterize methylation and hydroxymethylation changes in blood and brain associated with alcohol use disorder (AUD). First, we used epigenomic deconvolution to perform cell-type-specific methylome-wide association studies within subpopulations of granulocytes/T-cells/B-cells/monocytes in 1132 blood samples. Blood findings were then examined for overlap with AUD-related associations with methylation and hydroxymethylation in 50 human post-mortem brain samples. Follow-up analyses investigated if overlapping findings mediated AUD-associated transcription changes in the same brain samples. Lastly, we replicated our blood findings in an independent sample of 412 individuals and aimed to replicate published alcohol methylation findings using our results. Cell-type-specific analyses in blood identified methylome-wide significant associations in monocytes and T-cells. The monocyte findings were significantly enriched for AUD-related methylation and hydroxymethylation in brain. Hydroxymethylation in specific sites mediated AUD-associated transcription in the same brain samples. As part of the most comprehensive methylation study of AUD to date, this work involved the first cell-type-specific methylation study of AUD conducted in blood, identifying and replicating a finding in DLGAP1 that may be a blood-based biomarker of AUD. In this first study to consider the role of hydroxymethylation in AUD, we found evidence for a novel mechanism for cognitive deficits associated with AUD. Our results suggest promising new avenues for AUD research.
Im Bereich der Bilderkennung wurden in den vergangenen Jahren durch Deep Learning spektakuläre Fortschritte erzielt (Krizhevsky A, Sutskever I, Hinton GE. ImageNet classification with deep convolutional neural networks. In: Pereira F, Burges CJC, Bottou L, Weinberger KQ (Hrsg). Advances in Neural Information Processing Systems 25: Curran Associates, Inc; 2012, 1097–1105). Durch Neuronale Faltungsnetze (kurz: CNNs [convolutional neural networks]) können Straßenszenen in eine sogenannte semantische Segmentierung übersetzt werden (Long J, Shelhamer E, Darrell T. Fully convolutional networks for semantic segmentation. CoRR 2014; abs/1411.4038). Diese semantische Segmentierung bildet einen Baustein im Zusammenspiel mehrerer redundanter Systeme, welche das autonome Fahren ermöglichen sollen. Da neuronale Netze mithilfe von Trainingsdatensätzen angelernt werden, versagen diese oft bei dem Versuch unbekannte Objekte (also Objekte, welche nicht in den Trainingsdatensätzen vorhanden waren) zu erkennen.
OBJECTIVE:The impact of adolescent cannabis use is a pressing public health question owing to the high rates of use and links to negative outcomes. This study considered the association between problematic adolescent cannabis use and methylation.METHOD:Using an enrichment-based sequencing approach, a methylome-wide association study (MWAS) was performed of problematic adolescent cannabis use in 703 adolescent samples from the Great Smoky Mountain Study. Using epigenomic deconvolution, MWASs were performed for the main cell types in blood: granulocytes, T cells, B cells, and monocytes. Enrichment testing was conducted to establish overlap between cannabis-associated methylation differences and variants associated with negative mental health effects of adolescent cannabis use.RESULTS:Whole-blood analyses identified 45 significant CpGs, and cell type-specific analyses yielded 32 additional CpGs not identified in the whole-blood MWAS. Significant overlap was observed between the B-cell MWAS and genetic studies of education attainment and intelligence. Furthermore, the results from both T cells and monocytes overlapped with findings from an MWAS of psychosis conducted in brain tissue.CONCLUSION:In one of the first methylome-wide association studies of adolescent cannabis use, several methylation sites located in genes of importance for potentially relevant brain functions were identified. These findings resulted in several testable hypotheses by which cannabis-associated methylation can impact neurological development and inflammation response as well as potential mechanisms linking cannabis-associated methylation to potential downstream mental health effects.
images, sources and licenses for the Anomaly track of segmentmeifyoucan.com
BACKGROUND:Women are 1.5-3 times more likely to suffer from depression than men. This sex bias first emerges during puberty and then persists across the reproductive years. As the cause remains largely elusive, we performed a methylation-wide association study (MWAS) to generate novel hypotheses.METHODS:We assayed nearly all 28 million possible methylation sites in blood in 595 blood samples from 487 participants aged 9-17. MWASs were performed to identify methylation sites associated with increasing sex differences in depression symptoms as a function of pubertal stage. Epigenetic deconvolution was applied to perform analyses on a cell-type specific level.RESULTS:In monocytes, a substantial number of significant associations were detected after controlling the FDR at 0.05. These results could not be explained by plasma testosterone/estradiol or current/lifetime trauma. Our top results in monocytes were significantly enriched (ratio of 2.48) for genes in the top of a large genome-wide association study (GWAS) meta-analysis of depression and neurodevelopment-related Gene Ontology (GO) terms that remained significant after correcting for multiple testing. Focusing on our most robust findings (70 genes overlapping with the GWAS meta-analysis and the significant GO terms), we find genes coding for members of each of the major classes of axon guidance molecules (netrins, slits, semaphorins, ephrins, and cell adhesion molecules). Many of these genes were previously implicated in rodent studies of brain development and depression-like phenotypes, as well as human methylation, gene expression and GWAS studies.CONCLUSIONS:Our study suggests that the emergence of sex differences in depression may be related to the differential rewiring of brain circuits between boys and girls during puberty.
State-of-the-art semantic or instance segmentation deep neural networks (DNNs) are usually trained on a closed set of semantic classes. As such, they are ill-equipped to handle previously-unseen objects. However, detecting and localizing such objects is crucial for safety-critical applications such as perception for automated driving, especially if they appear on the road ahead. While some methods have tackled the tasks of anomalous or out-of-distribution object segmentation, progress remains slow, in large part due to the lack of solid benchmarks; existing datasets either consist of synthetic data, or suffer from label inconsistencies. In this paper, we bridge this gap by introducing the "SegmentMeIfYouCan" benchmark. Our benchmark addresses two tasks: Anomalous object segmentation, which considers any previously-unseen object category; and road obstacle segmentation, which focuses on any object on the road, may it be known or unknown.We provide two corresponding datasets together with a test suite performing an in-depth method analysis, considering both established pixel-wise performance metrics and recent component-wise ones, which are insensitive to object sizes. We empirically evaluate multiple state-of-the-art baseline methods, including several models specifically designed for anomaly / obstacle segmentation, on our datasets and on public ones, using our test suite.The anomaly and obstacle segmentation results show that our datasets contribute to the diversity and difficulty of both data landscapes.
Deep neural networks (DNNs) for the semantic segmentation of images are usually trained to operate on a pre-defined closed set of object classes. This is in contrast to the "open world" setting where DNNs are envisioned to be deployed to. From a functional safety point of view, the ability to detect so-called "out-of-distribution" (OoD) samples, i.e., objects outside of a DNN's semantic space, is crucial for many applications such as automated driving. A natural baseline approach to OoD detection is to threshold on the pixel-wise softmax entropy. We present a two-step procedure that significantly improves that approach. Firstly, we utilize samples from the COCO dataset as OoD proxy and introduce a second training objective to maximize the softmax entropy on these samples. Starting from pretrained semantic segmentation networks we re-train a number of DNNs on different in-distribution datasets and consistently observe improved OoD detection performance when evaluating on completely disjoint OoD datasets. Secondly, we perform a transparent post-processing step to discard false positive OoD samples by so-called "meta classification." To this end, we apply linear models to a set of hand-crafted metrics derived from the DNN's softmax probabilities. In our experiments we consistently observe a clear additional gain in OoD detection performance, cutting down the number of detection errors by 52% when comparing the best baseline with our results. We achieve this improvement sacrificing only marginally in original segmentation performance. Therefore, our method contributes to safer DNNs with more reliable overall system performance.
The majority of methylome-wide association studies (MWAS) have been performed using commercially available array-based technologies such as the Infinium Human Methylation 450K and the Infinium MethylationEPIC arrays (Illumina). While these arrays offer a convenient and relatively robust assessment of the probed sites they only allow interrogation of 2-4% of all CpG sites in the human genome. Methyl-binding domain sequencing (MBD-seq) is an alternative approach for MWAS that provides near-complete coverage of the methylome at similar costs as the array-based technologies. However, despite publication of multiple positive evaluations, the use of MBD-seq for MWAS is often fiercely criticized. Here we discuss key features of the method and debunk misconceptions using empirical data. We conclude that MBD-seq represents an excellent approach for large-scale MWAS and that increased utilization is likely to result in more discoveries, advance biological knowledge, and expedite the clinical translation of methylome-wide research findings.