The spectacular mineralogy and geochemistry of lamproites are inherited from their mantle source, which suggests heterogeneity of sub-continental lithospheric mantle (SCLM). Mantle heterogeneity may be a result of various metasomatic melt activities. The Gondwana basin lamprophyres and lamproites have some unique geochemical characters like higher abundances of TiO2, Nb, Hf, Zr and total REE which make them different from similar rock suites of other parts of the globe. These unique characters along with negative epsilon Nd(t) values (-2.88 to -3.52, this study) suggest that their mantle source might have been modified by multiple phases of metasomatism. The present research focuses on two major categories of lamproites-phlogopite lamproites from east Bokaro and olivine lamproites from west Bokaro coalfields. The phlogopite lamproites are relatively enriched in SiO2, K2O, Nd, Sm, Y, Zr and Hf compared to olivine lamproites, which are SiO2 poor, but enriched in MgO, Fe2O3T, CaO and Cr. The higher abundance of Nb and total REE along with high Zr/Hf and initial 87Sr/86Sr ratios in Bokaro lamproites suggest carbonatite metasomatism which was pervasive throughout the mantle peridotite. Higher abundance of SiO2, Al2O3, K2O, TiO2 and Rb in east Bokaro lamproites indicate metasomatism of the mantle source by alkaline-Ti-rich silicate melt, which was not all pervasive but occurred as metasomatic phlogopite-apatite-rutile veins. Besides these, evidence of relict ancient subduction-zone metasomatism is indicated by Ta-trough and Pb-crest in OIB-normalised diagram. It can be concluded that the mantle source for east Bokaro lamproites had greater inputs from phlogopite-apatite-rutile bearing metasomatic veins with lesser contributions of carbonatite-metasomatised peridotite wall-rock; whereas for west Bokaro lamproites the source inputs were dominantly derived from carbonatite-metasomatised peridotite wall-rock.
Fractured aquifer systems are discrete and often locally developed. So, exploration of groundwater in highly deformed crystalline terrains is challenging. In such terrains, shear-zone bifurcations, lithological contacts, and intersecting fracture-lineament networks are commonly considered as favourable sites for hydraulically significant fractured aquifer system. However presence of these structural features is not always sufficient for the formation of a potential aquifer. The present study, conducted in the Eastern Indian Precambrian metamorphic terrain in West Bengal, India, investigates the additional controlling factors responsible for development of viable fractured aquifer system and aims to improve understanding relevant to groundwater exploration in highly deformed crystalline rock terrain. Integrated geophysical surveys, including Electrical Resistivity Tomography (ERT), Vertical Electrical Sounding (VES), and Spontaneous Potential (S.P.) tomography, were employed for this purpose. Various types of lithological contacts are developed in brittle-ductile shear zone as such shear zone facilitates emplacement of different types of fluids along the brittle fractures in brittle ductile shear zone. The results indicate this variation in lithological contacts significantly influence the development of potential fractured aquifers. Lithology-dependent differential weathering plays a decisive role in determining hydraulically viable fractured aquifer system. Removal of weathered material from lithological contact is the pivotal factor for the development of viable aquifers in highly deformed terrain. Furthermore, SP method can serve as an effective tool for identifying locales associated with viable groundwater under comparable geological conditions.
We present GraspMolmo, a generalizable open-vocabulary task-oriented grasping (TOG) model. GraspMolmo predicts semantically appropriate, stable grasps conditioned on a natural language instruction and a single RGB-D frame. For instance, given "pour me some tea," GraspMolmo selects a grasp on a teapot handle rather than its body. Unlike prior TOG methods, which are limited by small datasets, simplistic language, and unrealistically simple scenes, GraspMolmo learns from PRISM, a novel large-scale synthetic dataset of 379k samples featuring complex environments and diverse, realistic task descriptions. We fine-tune the Molmo vision-language model on this data, enabling GraspMolmo to generalize to novel open-vocabulary instructions and objects. In challenging real-world evaluations, GraspMolmo achieves state-of-the-art results, with a 70% prediction success on complex tasks, compared to the 35% achieved by the next best alternative. GraspMolmo also successfully demonstrates the ability to predict semantically correct bimanual grasps zero-shot. We release our synthetic dataset, code, model, and benchmarks to accelerate research in task-semantic robotic manipulation, which, along with videos, are available at this URL.
Mineral exploration in regions of limited bedrock exposure depends on the excellence of the predictive model yielded from geophysical and geological studies. In this aspect, the accuracy of the positions, shapes, and size of the concealed ore bodies is important for later resource evaluation. Commonly used magnetic susceptibility surveys to explore buried magnetite deposits often fail to resolve the boundary between magnetite ore, and host rocks when the host rock contains ilmenite, and/or magnetite as an accessory mineral. Electrical Resistivity Imaging (ERI), and Self-Potential (SP), are better substitutes to resolve the issue and delineate the positions and shapes of the ore bodies in gabbroic host-rock in Purulia district, West Bengal, India. The concealed magnetite ore body showed a sharp decrease in electrical resistivity value in the 2D ERI study, and a significant negative SP value was concurrent with the inferred concealed magnetite bodies, compared to the gabbroic host rock. Hence, the combined result of 2D ERI and SP indicate analogous negative anomalies to the inferred magnetite ore bodies, verified by the surface geological information and mineralogical studies. Such geophysical anomalies could be combined with field data to reconstruct magnetite ore body modeling, providing a practical approach to prospect buried magnetite ore bodies in basic host rocks.
Reasoning goes beyond language; the real world requires reasoning about space, time, affordances, and much more that words alone cannot convey. Existing multimodal models exploring the potential of reasoning with images are brittle and do not scale. They rely on calling specialist tools, costly generation of images, or handcrafted reasoning data to switch between text and image thoughts. Instead, we offer a simpler alternative – Mull-Tokens – modality-agnostic latent tokens pre-trained to hold intermediate information in either image or text modalities to let the model think free-form towards the correct answer. We investigate best practices to train Mull-Tokens inspired by latent reasoning frameworks. We first train Mull-Tokens using supervision from interleaved text-image traces, and then fine-tune without any supervision by only using the final answers. Across four challenging spatial reasoning benchmarks involving tasks such as solving puzzles and taking different perspectives, we demonstrate that Mull-Tokens improve upon several baselines utilizing text-only reasoning or interleaved image-text reasoning, achieving a +3
Despite impressive high-level video comprehension, multimodal language models struggle with spatial reasoning across time and space. While current spatial training approaches rely on real-world video data, obtaining diverse footage with precise spatial annotations remains a bottleneck. To alleviate this bottleneck, we present SIMS-V – a systematic data-generation framework that leverages the privileged information of 3D simulators to create spatially-rich video training data for multimodal language models. Using this framework, we investigate which properties of simulated data drive effective real-world transfer through systematic ablations of question types, mixes, and scales. We identify a minimal set of three question categories (metric measurement, perspective-dependent reasoning, and temporal tracking) that prove most effective for developing transferable spatial intelligence, outperforming comprehensive coverage despite using fewer question types. These insights enable highly efficient training: our 7B-parameter video LLM fine-tuned on just 25K simulated examples outperforms the larger 72B baseline and achieves competitive performance with proprietary models on rigorous real-world spatial reasoning benchmarks. Our approach demonstrates robust generalization, maintaining performance on general video understanding while showing substantial improvements on embodied and real-world spatial tasks.
Reasoning about motion and space is a fundamental cognitive capability that is required by multiple real-world applications. While many studies highlight that large multimodal language models (MLMs) struggle to reason about space, they only focus on static spatial relationships, and not dynamic awareness of motion and space, i.e., reasoning about the effect of egocentric and object motions on spatial relationships. Manually annotating such object and camera movements is expensive. Hence, we introduce SAT, a simulated spatial aptitude training dataset utilizing 3D simulators, comprising both static and dynamic spatial reasoning across 175K question-answer (QA) pairs and 20K scenes. Complementing this, we also construct a small (150 image-QAs) yet challenging dynamic spatial test set using real-world images. Leveraging our SAT datasets and 6 existing static spatial benchmarks, we systematically investigate what improves both static and dynamic spatial awareness. Our results reveal that simulations are surprisingly effective at imparting spatial aptitude to MLMs that translate to real images. We show that perfect annotations in simulation are more effective than existing approaches of pseudo-annotating real images. For instance, SAT training improves a LLaVA-13B model by an average 11
While behavior cloning has recently emerged as a highly successful paradigm for autonomous driving, humans rarely learn to perform complex tasks, such as driving, via imitation or behavior cloning alone. In contrast, learning in humans often involves additional detailed guidance throughout the interactive learning process, i.e., where feedback, often via language, provides detailed information as to which part of their trial was performed incorrectly or suboptimally and why. Motivated by this observation, we introduce an efficient feedback-based framework for improving behavior-cloning-based training of sensorimotor driving agents. Our key insight is to leverage recent advances in Large Language Models (LLMs) to provide corrective fine-grained feedback regarding the underlying reason behind driving prediction failures. Moreover, our introduced network architecture is efficient, enabling the first sensorimotor end-to-end training and evaluation of LLM-based driving models. The resulting agent achieves state-of-the-art performance in open-loop evaluation on nuScenes, outperforming prior state-of-the-art by over 8.1% and 57.1% in accuracy and collision rate, respectively. In CARLA, our camera-based agent improves by 16.6% in driving score over prior LIDAR-based approaches.
An east–west‐trending medium‐grained mafic sill containing co‐genetic Fe–Ti oxide ore lenses is found disposed within granite gneisses around Saltora‐Mejia area in the eastern part of the Chotanagpur Granite Gneissic Complex (CGGC) of eastern India. CGGC is considered as a Proterozoic mobile belt as it witnessed multiple phases of deformation and high‐ grade metamorphism during 1.8–0.8 Ga. Occurrence of such Fe–Ti oxide ore‐bearing mafic sill is unique in the entire CGGC which is a vast Proterozoic orogenic belt and has witnessed many phases of voluminous mafic and felsic magmatisms. The mafic rock is of gabbronorite composition which contains plagioclase, clinopyroxene, orthopyroxene as major constituent primary minerals and amphibole as late magmatic mineral. The rock shows sub‐ophitic, intergranular, mosaic and poikilitic texture (defined by larger pargasitic grain). The gabbronorite shows iron enriched tholeiitic character, low Mg#, low abundances of Ni and Cr, slight enrichment in LILE, LREE and slight depletion in HFSE like Nb and Ti. The computed melt in equilibrium with the studied gabbronorite shows transitional orogenic to anorogenic, within‐plate and E‐MORB‐like geochemical character. In this study, the U–Pb zircon crystallization age (~960 Ma) of the Saltora‐Mejia gabbronorite is reported for the first time which coincides with the late tectonic stage of the most pervasive orogenic activity in the CGGC around 1.2–0.9 Ga. Transitional orogenic to anorogenic geochemical character, late tectonic evolution and other field and laboratory evidences together suggest evolution of the Saltora‐Mejia gabbronorite sill in a late tectonic extensional environment which might have been facilitated by delamination of a subducted plate and upwelling of asthenospheric mantle during the waning stage of a major orogeny in the CGGC.
We propose a novel VQA dataset, BloomVQA, to facilitate comprehensive evaluation of large vision-language models on comprehension tasks. Unlike current benchmarks that often focus on fact-based memorization and simple reasoning tasks without theoretical grounding, we collect multiple-choice samples based on picture stories that reflect different levels of comprehension, as laid out in Bloom's Taxonomy, a classic framework for learning assessment widely adopted in education research. Our data maps to a novel hierarchical graph representation which enables automatic data augmentation and novel measures characterizing model consistency. We perform graded evaluation and reliability analysis on recent multi-modal models. In comparison to lowlevel tasks, we observe decreased performance on tasks requiring advanced comprehension and cognitive skills with up to 38.0% drop in VQA accuracy. In comparison to earlier models, GPT-4V demonstrates improved accuracy over all comprehension levels and shows a tendency of bypassing visual inputs especially for higher-level tasks. Current models also show consistency patterns misaligned with human comprehension in various scenarios, demonstrating the need for improvement based on theoretically-grounded criteria. The dataset can be accessed at https://huggingface. co/datasets/ygong/BloomVQA.
Professional artists, photographers, and other visual content creators use object relighting to establish their photo's desired effect. Unfortunately, manual tools that allow relighting have a steep learning curve and are difficult to master. Although generative editing methods now enable some forms of image editing, relighting is still beyond today's capabilities; existing methods struggle to keep other aspects of the image -- colors, shapes, and textures -- consistent after the edit. We propose Lasagna, a method that enables intuitive text-guided relighting control. Lasagna learns a lighting prior by using score distillation sampling to distill the prior of a diffusion model, which has been finetuned on synthetic relighting data. To train Lasagna, we curate a new synthetic dataset ReLiT, which contains 3D object assets re-lit from multiple light source locations. Despite training on synthetic images, quantitative results show that Lasagna relights real-world images while preserving other aspects of the input image, outperforming state-of-the-art text-guided image editing methods. Lasagna enables realistic and controlled results on natural images and digital art pieces and is preferred by humans over other methods in over 91% of cases. Finally, we demonstrate the versatility of our learning objective by extending it to allow colorization, another form of image editing.
We propose a self-supervised approach for learning to perform audio source separation in videos based on natu-ral language queries, using only unlabeled video and au-dio pairs as training data. A key challenge in this task is learning to associate the linguistic description of a sound-emitting object to its visual features and the corresponding components of the audio waveform, all without access to annotations during training. To overcome this challenge, we adapt off-the-shelf vision-language foundation models to provide pseudo-target supervision via two novel loss functions and encourage a stronger alignment between the audio, visual and natural language modalities. During inference, our approach can separate sounds given text, video and audio input, or given text and audio input alone. We demonstrate the effectiveness of our self-supervised approach on three audio-visual separation datasets, including MUSIC, SOLOS and AudioSet, where we outperform state-of-the-art strongly supervised approaches despite not using object detectors or text labels during training. Our project page including publicly available code can be found at https://cs-people.bu.edu/rxtan/projectsNAST.
Compositional reasoning is a hallmark of human visual intelligence. Yet, despite the size of large vision-language models, they struggle to represent simple compositions by combining objects with their attributes. To measure this lack of compositional capability, we design Cola, a text-to-image retrieval benchmark to Compose Objects Localized with Attributes. To solve Cola, a model must retrieve images with the correct configuration of attributes and objects and avoid choosing a distractor image with the same objects and attributes but in the wrong configuration. Cola contains about 1.2k composed queries of 168 objects and 197 attributes on around 30K images. Our human evaluation finds that Cola is 83.33% accurate, similar to contemporary compositionality benchmarks. Using Cola as a testbed, we explore empirical modeling designs to adapt pre-trained vision-language models to reason compositionally. We explore 6 adaptation strategies on 2 seminal vision-language models, using compositionality-centric test benchmarks - Cola and CREPE. We find the optimal adaptation strategy is to train a multi-modal attention layer that jointly attends over the frozen pre-trained image and language features. Surprisingly, training multimodal layers on CLIP performs better than tuning a larger FLAVA model with already pre-trained multimodal layers. Furthermore, our adaptation strategy improves CLIP and FLAVA to comparable levels, suggesting that training multimodal layers using contrastive attribute-object data is key, as opposed to using them pre-trained. Lastly, we show that Cola is harder than a closely related contemporary benchmark, CREPE, since simpler fine-tuning strategies without multimodal layers suffice on CREPE but not on Cola. However, we still see a significant gap between our best adaptation and human accuracy, suggesting considerable room for further research.
Existing emotion prediction benchmarks contain coarse emotion labels which do not consider the diversity of emotions that an image and text can elicit in humans due to various reasons. Learning diverse reactions to multimodal content is important as intelligent machines take a central role in generating and delivering content to society. To address this gap, we propose Socratis, a societal reactions benchmark, where each image-caption (IC) pair is annotated with multiple emotions and the reasons for feeling them. Socratis contains 18K free-form reactions for 980 emotions on 2075 image-caption pairs from 5 widely-read news and image-caption (IC) datasets. We benchmark the capability of state-of-the-art multimodal large language models to generate the reasons for feeling an emotion given an IC pair. Based on a preliminary human study, we observe that humans prefer human-written reasons over 2 times more often than machine-generated ones. This shows our task is harder than standard generation tasks because it starkly contrasts recent findings where humans cannot tell apart machine vs human-written news articles, for instance. We further see that current captioning metrics based on large vision-language models also fail to correlate with human preferences. We hope that these findings and our benchmark will inspire further research on training emotionally aware models.
Introduction: Identifying the Paravertebral Space (PVS) by its anatomical landmarks is associated with high failure rates and complications. With the advent of Ultrasonography (USG), failure rate has decreased leading to an increased interest in performing USG guided Thoracic Paravertebral Block (TPVB). Aim: To assess the efficacy and safety of ultrasound guided TPVB and its comparison with the landmark-based technique, in patients undergoing elective unilateral breast surgery. Materials and Methods: This cross-sectional study was carried out at Command Hospital, Pune from July 2014 to December 2015, on females between 18-70 years, accepted in American Society of Anaesthesiology (ASA) I-III for unilateral breast surgeries. Patients were divided into two groups with 40 subjects in each group. Group A subjects were treated with anatomical landmark technique and group B subjects with USG guided technique. The p-value <0.05 was considered to be statistically significant. Results: Demographic parameters (age, height, weight and Body Mass Index (BMI) and the scheduled surgery were comparable in between the groups. In group A, success rate of the block was 82.5%, compared to 95% in group B (p-value >0.05 using Fisher’s-exact test). Mean (SD) time taken for performing the block in group A was 371.10 (10.37) seconds while it was 613.73 (37.15) seconds in group B (p-value <0.05 by two independent sample t-tests). No statistically significant difference was seen in haemodynamic parameters, except for the Heart Rate (HR) at 70, 80, 90 minutes after administering the block and at the end of surgery. Correlation analysis for quantitative variables with PVS depth (dependent variable), measured sonologically, showed very good linear correlation of PVS depth with weight (Pearson correlation coefficient, r=0.819, p-value <0.001). BMI (r=0.884; p-value <0.001). Conclusion: The success rate is higher with ultrasound guided TPVB compared to the landmark technique though statistically insignificant. But it is recommended to use ultrasound-guided TPVB for advantages such as lesser requirement of opioid supplementation, real time visualisation of the spread of drugs in PVS with lesser complication rates.
Large-scale pre-trained vision-and-language 001 (V+L) transformers have propelled the state 002 of the art (SOTA) on Visual Question Answer-003 ing (VQA) task. Despite impressive perfor-004 mance on the standard VQA benchmark, it re-005 mains unclear how robust these models are. To 006 investigate, we conduct a host of evaluations 007 over 4 different types of robust VQA datasets: 008 ( i ) Linguistic Variation; ( ii ) Logical Reason-009 ing; ( iii ) Visual Content Manipulation; and 010 ( iv ) Answer Distribution Shift. Experiments 011 show that pre-trained V+L models already ex-012 hibit better robustness than many task-specific 013 SOTA methods via standard model finetun-014 ing. To further enhance model robustness, we 015 propose M ANGO , a generic and efficient ap-016 proach that learns a M ultimodal A dversarial 017 N oise G enerat O r in the embedding space to 018 fool V+L models. Differing from previous 019 studies focused on one specific type of robust-020 ness, M ANGO is agnostic to robustness types, 021 and enables universal performance lift for both 022 task-specific and pre-trained models over di-023 verse robust VQA datasets designed to evaluate 024 broad aspects of robustness. Comprehensive 025 experiments demonstrate that M ANGO outper-026 forms previous task-specific SOTAs on 7 out 027 of 9 robustness benchmarks. 028
Bhanjada Bet igneous rock lies to the east of Pachham Island and west of Khadir Island, well within the Great Rann of Kutch, Gujarat, western India. It is a small isolated hillock made up largely of phonolite with patches of trachyte and traversed by a mafic dyke. The phonolite is composed dominantly of sanidine, nepheline, aegirine, Ti-amphibole and glass. Sanidine occurs as phenocryst and nepheline as microphenocryst. Groundmass is composed of smaller alkali feldspar, aegirine, Ti-amphibole and glass. The lath-shaped feldspar in the groundmass defines excellent flow texture. Major element chemical composition indicates occurrence of two groups of phonolite-low silica (around 53%) and high silica (around 58%). The phonolite has higher abundance of MgO (2.21–3.5%) over FeO (2.43–2.6%) and very high Mg# (65–69) suggesting its primitive nature. Trace element abundances suggest an overall enrichment in LILE and selected HFSE, highly fractionated LREE pattern. High Mg# of the Bhanjada bet phonolite along with low values of Zr and Hf compared to phonolites of basanite phonolite suite where phonolite is an evolved member, is noted. The lower Zr and Hf values hint towards a less evolved primary phonolitic magma rather than phonolite as late differentiates of basanitic–phonolite suite. Experimental studies indicate low degree partial melting of a metasomatised lherzolite mantle at a pressure around 1–1.5 GPa can produce primary phonolite magma. Mantle xenoliths, found in alkali basalt of Kutch basin, show evidence of carbonatite metasomatism of lithospheric mantle. Bhanjada Bet phonolites are associated with layered mafic complex of Nir Wandh and other magmatic rocks of Pachham Island. Ages of Nir Wandh magmatic rocks and magmatic rocks of Pachham Island coincide with ages of early Deccan magmatism. Age of phonolite is not yet known. From field association and petro-mineralogy, we propose that Bhanjada Bet phonolite represents alkali magmatism of early Deccan age in Western Deccan Province.
The effects of deformation and concomitant high‐grade metamorphism during the Cenozoic period in the active Himalayan orogen hinder the reconstruction of the original geological framework and thus the evolution history. The Shillong Group in the Shillong Plateau, northeastern India, rests unconformably above the Himalayan crystalline basement gneiss, and is divided in two formations, the Lower Metapelite Formation (LMF) and the Upper Quartzite Formation (UQF). There are three horizons of conglomerate: (a) Basal, between the basement and the Shillong Group, (b) an inter‐formational conglomerate, between the LMF and UQF, and (c) an intra‐formational conglomerate (within the UQF). The clasts of the latter two conglomerates are suitable strain markers. Principal compressional structures in the Shillong Group were developed in two successive deformation episodes during the Meso‐ to Neo‐Proterozoic period. The earlier episode was of progressive general shear deformation, while the later one possibly occurred in a transpressive mode similar to that of a fold‐thrust belt. Remarkably, both episodes underwent NW–SE regional compression. The inter‐formational conglomerate represents a flattening type strain of relatively lesser magnitude ( R f – ϕ X : Y : Z = 1.35:1:0.76), in contrast to the constriction type strain of relatively higher magnitude ( R f – ϕ X : Y : Z = 2.42:1:0.44) of the intra‐formational conglomerate. Field relations of the Shillong Group with the deformed and undeformed granite plutons suggest that the earlier deformation episode possibly took place during Rodinia assembly, due to India–Antarctica collision at 1,100 Ma, while the later one happened in a time frame close to 500 Ma, during the amalgamation of Eastern Gondwana.
The present study deals with the petrogenesis and age implication of anorogenic peralkaline granitoid rocks exposed along the North Puruliya Shear Zone (NPSZ) in Jhalda area of Puruliya district, West Bengal. Alkali granite consists of quartz, alkali feldspar, aegirine, riebeckite, arfvedsonite and biotite. These granitoid rocks have a high range of silica, very high total alkali content and are poor in CaO, Al 2 O 3 , FeO and MgO content. Geochemically, they are ferroan, alkalic, reduced and peralkaline granitoid rocks and have many similarities with A-type granites. Crystallisation temperatures of these granitoid rocks are greater than 900°C. U–Pb isotopic ages of zircon indicate a major age cluster ~966.7 ± 7.0 Ma. The oldest lower crustal rocks in and around Jhalda area are charnockite, khondalite, garnetiferous granite gneiss, which might have acted as source rocks. Trace element model indicates that a moderate degree partial melting (5–20%) of charnockite + khondalite source rock followed by ~30% fractional crystallisation of plagioclase feldspar is responsible to generate parent magma of alkali granite. Similar and overlapping crystallisation ages of 966.7 ± 7.0 Ma of the per-alkaline anorogenic/post-orogenic granites of present study with already reported orogenic I-type granites from Jhalda and S-type granites from nearby Raghunathpur area (age 1000 Ma) may indicate origin and emplacement of post-orogenic granites of Jhalda during orogeny–anorogeny transition at the time of waning stage of orogenic activity. Mantle upwelling in late to post-orogenic stage provides additional heat to initiate partial melting of lower crustal source rocks.