The timing of incision of the Yellow River at Sanmenxia Gorge has long been a key question in understanding the formation of the modern Yellow River. Despite extensive geomorphological, provenance, and tectonic research, debates persist regarding the timing of transition from an endorheic to exoreic river system. To address this issue, we conducted paleomagnetic dating, grain-size analysis, and magnetic susceptibility measurements on a 200-mlong core (GT3) from the southern Weihe Basin in the middle reaches of the Yellow River. Our results revealed that the sedimentary environment in this area transitioned through four stages: (1) a shallow lacustrine phase (similar to 3.69-2.58 Ma), (2) fluvial facies (similar to 2.58-1.8 Ma), (3) alternating fluvial and aeolian processes (similar to 1.8-1.3 Ma), and (4) an irreversible drought phase beginning at similar to 1.3 Ma. Our findings suggest that the modern drainage pattern of the middle reaches of the Yellow River was essentially established in the early Pleistocene. Tectonic uplift or differential subsidence across the study region played a pivotal role in the sedimentary evolution of the Weihe Basin. This study provides novel insights into the formation of the middle reaches of the Yellow River and its driving mechanisms, particularly highlighting the connectivity dynamics of Sanmenxia Gorge as the river shifted eastwards establishing an outflow into the Yellow Sea.
Vision-language models (VLMs) are widely assumed to exhibit in-context learning (ICL), a property similar to that of their language-only counterparts. While recent work suggests VLMs can perform multimodal ICL (MM-ICL), studies show they often rely on shallow heuristics -- such as copying or majority voting -- rather than true task understanding. We revisit this assumption by evaluating VLMs under distribution shifts, where support examples come from a dataset different from the query. Surprisingly, performance often degrades with more demonstrations, and models tend to copy answers rather than learn from them. To investigate further, we propose a new MM-ICL with Reasoning pipeline that augments each demonstration with a generated rationale alongside the answer. We conduct extensive and comprehensive experiments on both perception- and reasoning-required datasets with open-source VLMs ranging from 3B to 72B and proprietary models such as Gemini 2.0. We conduct controlled studies varying shot count, retrieval method, rationale quality, and distribution. Our results show limited performance sensitivity across these factors, suggesting that current VLMs do not effectively utilize demonstration-level information as intended in MM-ICL.
Analyzing multivariate time series is important in many domains. However, it has been difficult to learn robust and generalizable representations within multivariate datasets due to complex inter-channel relationships and dynamic shifts. In this paper, we introduce a novel approach for learning spatiotemporal structure and using it to improve the application of transformers to timeseries datasets. Our framework learns a set of group tokens, and builds an instance-specific group embedding (GE) layer that assigns input tokens to a small number of group tokens to incorporate structure into learning. We then introduce a novel architecture, Group-Aware transFormer (GAFormer), which incorporates both spatial and temporal group embeddings to achieve state-of-the-art performance on a number of time-series classification and regression tasks. In evaluations on a number of diverse timeseries datasets, we show that GE on its own can provide a nice enhancement to a number of backbones, and that by coupling spatial and temporal group embeddings, the GAFormer can outperform the existing baselines. Finally, we show how our approach discerns latent structures in data even without information about the spatial ordering of channels, and yields a more interpretable decomposition of spatial and temporal structure underlying complex multivariate datasets.
Despite significant advances in deep learning, models often struggle to generalize well to new, unseen domains, especially when training data is limited. To address this challenge, we propose a novel approach for distribution-aware latent augmentation that leverages the relationships across samples to guide the augmentation procedure. Our approach first degrades the samples stochastically in the latent space, mapping them to augmented labels, and then restores the samples from their corrupted versions during training. This process confuses the classifier in the degradation step and restores the overall class distribution of the original samples, promoting diverse intra-class/crossdomain variability. We extensively evaluate our approach on a diverse set of datasets and tasks, including domain generalization benchmarks and medical imaging datasets with strong domain shift, where we show our approach achieves significant improvements over existing methods for latent space augmentation. We further show that our method can be flexibly adapted to long-tail recognition tasks, demonstrating its versatility in building more generalizable models. https://github.com/nerdslab/LatentDR.
Time series data are inherently functions of time, yet current transformers often learn time series by modeling them as mere concatenations of time periods, overlooking their functional properties. In this work, we propose a novel objective for transformers that learn time series by re-interpreting them as temporal functions. We build an alternative sequence of time series by constructing degradation operators of different intensity in the functional space, creating augmented variants of the original sample that are abstracted or simplified to different degrees. Based on the new set of generated sequence, we train an autoregressive transformer that progressively recovers the original sample from the most simplified variant. Analogous to the next word prediction task in languages that learns narratives by connecting different words, our autoregressive transformer aims to learn the Narratives of Time Series (NoTS) by connecting different functions in time. Theoretically, we justify the construction of the alternative sequence through its advantages in approximating functions. When learning time series data with transformers, constructing sequences of temporal functions allows for a broader class of approximable functions (e.g., differentiation) compared to sequences of time periods, leading to a 26% performance improvement in synthetic feature regression experiments. Experimentally, we validate NoTS in 3 different tasks across 22 real-world datasets, where we show that NoTS significantly outperforms other pre-training methods by up to 6%. Additionally, combining NoTS on top of existing transformer architectures can consistently boost the performance. Our results demonstrate the potential of NoTS as a general-purpose dynamic learner, offering a viable alternative for developing foundation models for time series analysis.
We analyzed the trace and rare earth element contents of the desert sands and loess deposits in the Junggar Basin. Combined with previously published data, and using principal component analysis, our results provide insights into the sand provenances of the Gurbantunggut Desert in eastern Central Asia (CA); the genetic links between deserts and loess deposits; and the specific sources in CA for the aeolian dust in North Pacific Ocean sediments. The results also demonstrate the spatial heterogeneity of the geochemistry of sand across the Gurbantunggut Desert. The desert sands in the northern and western parts of this desert are mainly derived from the Altai and Junggar mountains, respectively, as supported by the north-south directions of sand fluxes. However, Beitashan Mountain makes a negligible contribution due to the lack of fluvial transport and westward sand fluxes. However, more sediment samples need to be collected to confirm the contribution of a “Tianshan” Mountains source. Our findings also indicate that the Gurbantunggut Desert did not contribute significantly to loess accumulation on the northern slopes of the Tianshan Mountains, in accordance with the weak genetic relationships between the loess deposits and deserts in western CA. We attribute this to the limited ability of the CA deserts to produce and supply dust-sized particles. Additionally, using the Metropolis-Hastings sampling approach, we found that the Gurbantunggut Desert is not the source region in CA for the fine dust particles in North Pacific Ocean sediments. Overall, our results contribute to a deeper understanding of the aeolian systems in CA, and they elucidate their impacts on the dust cycle at a hemispheric scale.
Understanding the climatic evolution in Central Asia (CA) and its drivers is crucial for informed decision-making and predicting global changes due to its significant contribution as a global dust source. Unfortunately, our current understanding of the pre-Holocene precipitation patterns in CA is lacking due to the limited availability of reliable proxy indicators, and our knowledge of future precipitation projections in the region, based on paleoclimatic dynamics, is also quite limited. In this study, we analyzed variations in carbonate and dolomite contents of a loess section in the Ili Basin, northern CA, to reveal precipitation changes during the last glacial period. The results showed that changes in carbonate minerals were mainly influenced by the source material supply, driven by reduced precipitation and eluviation during glacial period. We thereby established a precipitation index by removing the influence of provenance signals from the dolomite records. The index indicated lower precipitation during mid-marine isotope stage (MIS) 3 compared to MIS2, likely due to meridional shifts and intensity changes of the westerlies caused by changes in precession and obliquity, with precession playing a major role. Through the comparison of the precipitation index with the delta 18O records of the Greenland ice core on a millennial timescale, it was observed that the precipitation in northern CA exhibited a positive correlation with the North Atlantic Oscillation (NAO) mode due to migration of the westerlies. By leveraging our understanding of orbital- and millennial-scale precipitation patterns, we utilized the random forest (RF) regression model and the autoregressive integrated moving average model to forecast precipitation changes for the upcoming 5000-10,000 years. The results indicated a variable pattern marked by a general upward trend, suggesting the possibility of favorable development of agricultural-based economies in the Ili River Valley. People should realize that some integrated measures are designed to improve resilience of agricultural sector in the region and enhance its capacity to adapt to challenges posed by climate change. However, more extensive research is necessary to verify these results through thorough examination and comparisons of loess sections in our research location.
Despite the success of contrastive learning (CL) in vision and language, its theoretical foundations and mechanisms for building representations remain poorly understood. In this work, we build connections between noise contrastive estimation losses widely used in CL and distribution alignment with entropic optimal transport (OT). This connection allows us to develop a family of different losses and multistep iterative variants for existing CL methods. Intuitively, by using more information from the distribution of latents, our approach allows a more distribution-aware manipulation of the relationships within augmented sample sets. We provide theoretical insights and experimental evidence demonstrating the benefits of our approach for generalized contrastive alignment. Through this framework, it is possible to leverage tools in OT to build unbalanced losses to handle noisy views and customize the representation space by changing the constraints on alignment. By reframing contrastive learning as an alignment problem and leveraging existing optimization tools for OT, our work provides new insights and connections between different self-supervised learning models in addition to new tools that can be more easily adapted to incorporate domain knowledge into learning.
X-ray diffraction (XRD) analysis, as one of the most powerful methods, has been widely used to identify and quantify minerals in earth science. How to improve the precision of mineral quantitative analysis is still a hot topic. To date, several quantitative methods have been proposed for different purposes and accompanied by diverse software. In this study, three quantitative mineral analysis methods, including the reference intensity ratio (RIR), Rietveld, and full pattern summation (FPS) methods, are compared and evaluated to systematically investigate their accuracy and applicability. The results show that the analytical accuracy of these methods is basically consistent for mixtures free from clay minerals. However, there are significant differences in accuracy for clay-mineral-containing samples. In comparison, it seems that the FPS method has wide applicability, which is more appropriate for sediments. The Rietveld method has been shown to be capable of quantifying complicated non-clay samples with a high analytical accuracy; nevertheless, most conventional Rietveld software fails to accurately quantify phases with a disordered or unknown structure. The RIR method represents a handy approach but with lower analytical accuracy. Overall, the present results are expected to provide a potentially important reference for the quantitative analysis of minerals in sediments.
There are multiple scales of abstraction from which we can describe the same image, depending on whether we are focusing on fine-grained details or a more global attribute of the image. In brain mapping, learning to automatically parse images to build representations of both small-scale features (e.g., the presence of cells or blood vessels) and global properties of an image (e.g., which brain region the image comes from) is a crucial and open challenge. However, most existing datasets and benchmarks for neuroanatomy consider only a single downstream task at a time. To bridge this gap, we introduce a new dataset, annotations, and multiple downstream tasks that provide diverse ways to readout information about brain structure and architecture from the same image. Our multi-task neuroimaging benchmark (MTNeuro) is built on volumetric, micrometer-resolution X-ray microtomography images spanning a large thalamocortical section of mouse brain, encompassing multiple cortical and subcortical regions. We generated a number of different prediction challenges and evaluated several supervised and self-supervised models for brain-region prediction and pixel-level semantic segmentation of microstructures. Our experiments not only highlight the rich heterogeneity of this dataset, but also provide insights into how self-supervised approaches can be used to learn representations that capture multiple attributes of a single image and perform well on a variety of downstream tasks. Datasets, code, and pre-trained baseline models are provided at: https://mtneuro.github.io/ .
Complex time-varying systems are often studied by abstracting away from the dynamics of individual components to build a model of the population-level dynamics from the start. However, when building a population-level description, it can be easy to lose sight of each individual and how they contribute to the larger picture. In this paper, we present a novel transformer architecture for learning from time-varying data that builds descriptions of both the individual as well as the collective population dynamics. Rather than combining all of our data into our model at the onset, we develop a separable architecture that operates on individual time-series first before passing them forward; this induces a permutation-invariance property and can be used to transfer across systems of different size and order. After demonstrating that our model can be applied to successfully recover complex interactions and dynamics in many-body systems, we apply our approach to populations of neurons in the nervous system. On neural activity datasets, we show that our model not only yields robust decoding performance, but also provides impressive performance in transfer across recordings of different animals without any neuron-level correspondence. By enabling flexible pre-training that can be transferred to neural recordings of different size and order, our work provides a first step towards creating a foundation model for neural decoding.
Lip reading is the task of recognizing speech content by analyzing movements in the lip region when people are speaking. Based on the continuity in adjacent frames in the speaking process, and the consistency in motion patterns among different people when they pronounce the same phoneme, we model lip movements as a sequence of apparent deformations in the lip region during the speaking process. Specifically, we introduce a Deformation Flow Network (DFN) to learn the deformation flow between adjacent frames, which directly captures the motion information within the lip region. The learned deformation flow is then combined with the original grayscale frames with a two-stream network to perform lip reading. To make the two streams learn from each other in the learning process, we introduce a bidirectional knowledge distillation loss to train the two branches jointly. Owing to the complementary cues provided by different branches, the two-stream network shows substantial improvement over using either single branch. A thorough experimental evaluation on two large-scale lip reading benchmarks is presented with detailed analysis. The results accord with our motivation, and show that our method achieves state-of-the-art or comparable performance on these two challenging datasets.
Recent advances in deep learning have heightened interest among researchers in the field of visual speech recognition (VSR). Currently, most existing methods equate VSR with automatic lip reading, which attempts to recognise speech by analysing lip motion. However, human experience and psychological studies suggest that we do not always fix our gaze at each other's lips during a face-to-face conversation, but rather scan the whole face repetitively. This inspires us to revisit a fundamental yet somehow overlooked problem: can VSR models benefit from reading extraoral facial regions, i.e. beyond the lips? In this paper, we perform a comprehensive study on the evaluation of the effects of different facial regions with state-of-the-art VSR models, including the mouth, the whole face, the upper face, and even the cheeks. Experiments are conducted on both word-level and sentence-level benchmarks with different characteristics. We find that despite the complex variations of the data, incorporating information from extraoral facial regions, even the upper face, consistently benefits VSR performance. Furthermore, we introduce a simple yet effective method based on Cutout to learn more discriminative features for face-based VSR, hoping to maximise the utility of information encoded in different facial regions. Our experiments show obvious improvements over existing state-of-the-art methods that use only the lip region as inputs, a result we believe would probably provide the VSR community with some new and exciting insights.
Large-scale datasets have successively proven their fundamental importance in several research fields, especially for early progress in some emerging topics. In this paper, we focus on the problem of visual speech recognition, also known as lipreading, which has received increasing interest in recent years. We present a naturally-distributed large-scale benchmark for lip reading in the wild, named LRW-1000, which contains 1,000 classes with 718,018 samples from more than 2,000 individual speakers. Each class corresponds to the syllables of a Mandarin word composed of one or several Chinese characters. To the best of our knowledge, it is currently the largest word-level lipreading dataset and also the only public large-scale Mandarin lipreading dataset. This dataset aims at covering a "natural" variability over different speech modes and imaging conditions to incorporate challenges encountered in practical applications. It has shown a large variation in this benchmark in several aspects, including the number of samples in each class, video resolution, lighting conditions, and speakers' attributes such as pose, age, gender, and make-up. Besides providing a detailed description of the dataset and its collection pipeline, we evaluate several typical popular lipreading methods and perform a thorough analysis of the results from several aspects. The results demonstrate the consistency and challenges of our dataset, which may open up some new promising directions for future work.
This report describes the approach underlying our submission to the active speaker detection task (task B-2) of ActivityNet Challenge 2019. We introduce a new audio-visual model which builds upon a 3D-ResNet18 visual model pretrained for lipreading and a VGG-M acoustic model pretrained for audio-to-video synchronization. The model is trained with two losses in a multi-task learning fashion: a contrastive loss to enforce matching between audio and video features for active speakers, and a regular crossentropy loss to obtain speaker / non-speaker labels. This model obtains 84.0% mAP on the validation set of AVAActiveSpeaker. Experimental results showcase the pretrained embeddings’ abilities to transfer across tasks and data formats, as well as the advantage of the proposed multi-task learning strategy.
Visual speech recognition is the task to decode the speech content from a video based on visual information, especially the movements of lips. It is also referenced as lipreading. Motivated by two problems existing in lipreading, words with similar pronunciation and the variation of word duration, we propose a novel 3D Feature Pyramid Attention (3D-FPA) module to jointly improve the representation power of features in both the spatial and temporal domains. Specifically, the input features are downsampled for 3 times in both the spatial and temporal dimensions to construct spatiotemporal feature pyramids. Then high-level features are upsampled and combined with low-level features, finally generating a pixel-level soft attention mask to be multiplied with the input features.It enhances the discriminative power of features and exploits the temporal multi-scale information while decoding the visual speeches. Also, this module provides a new method to construct and utilize temporal pyramid structures in video analysis tasks. The field of temporal featrue pyramids are still under exploring compared to the plentiful works on spatial feature pyramids for image analysis tasks. To validate the effectiveness and adaptability of our proposed module, we embed the module in a sentence-level lipreading model, LipNet, with the result of 3.6% absolute decrease in word error rate, and a word-level model, with the result of 1.4% absolute improvement in accuracy.
Edge detection is a classic problem in the field of image processing, which lays foundations for other tasks such as image segmentation. Conventionally, this operation is performed using gradient operators such as the Roberts or Sobel operator, which can discover local changes in intensity levels. These operators, however, perform poorly on low contrast images. In this paper, we propose an edge detector architecture for color images based on fuzzy theory and the Sobel operator. First, the R, G and B channels are extracted from an image and enhanced using fuzzy methods, in order to suppress noise and improve the contrast between the background and the objects. The Sobel operator is then applied to each of the channels, which are finally combined into an edge map of the origin image. Experimental results obtained through an FPGA-based implementation have proved the proposed method effective.