Large Language Models (LLMs) exhibit In-Context Learning (ICL), which enables the model to perform new tasks conditioning only on the examples provided in the context without updating the model’s weights. While ICL offers fast adaptation across natural language tasks and domains, its emergence is less straightforward for modalities beyond text. In this work, we systematically uncover properties present in LLMs that support the emergence of ICL for autoregressive models and various modalities by promoting the learning of the mechanisms needed for ICL. We identify exact token repetitions in the training data sequences as an important factor for ICL. Such repetitions further improve stability and reduce transiency in ICL performance. We analyse in detail the training dynamics of such data sequences and explain how token repetitions enhance the ICL learning mechanisms. Moreover, we emphasise the importance of the training task difficulty for the emergence of ICL. Finally, by applying our novel insights on ICL emergence, we unlock ICL capabilities across various visual datasets used for few-shot classification, and confirm the generalisability of our insights to much harder real-world examples of large-scale object classification, and a more challenging EEG classification task. Code is available at https://github.com/jelenab98/unlocking_icl
Large Language Models (LLMs) exhibit In-Context Learning (ICL), which enables the model to perform new tasks conditioning only on the examples provided in the context without updating the model’s weights. While ICL offers fast adaptation across natural language tasks and domains, its emergence is less straightforward for modalities beyond text. In this work, we systematically uncover properties present in LLMs that support the emergence of ICL for autoregressive models and various modalities by promoting the learning of the needed mechanisms for ICL. We identify exact token repetitions in the training data sequences as an important factor for ICL. Such repetitions further improve stability and reduce transiency in ICL performance. Moreover, we emphasise the significance of training task difficulty for the emergence of ICL. Finally, by applying our novel insights on ICL emergence, we unlock ICL capabilities for various visual datasets and a more challenging EEG classification task. Code is available at https://github.com/jelenab98/unlocking_icl .
Current world models operate at a single level of abstraction, with most prioritizing perceptual fidelity while lacking the spatial reasoning and semantic understanding required for real-world downstream tasks. We present a hierarchical driving world model that factorizes future prediction across two levels operating at distinct temporal and abstraction scales: a high-level predictor that forecasts coarse scene structure over extended temporal horizons, and a low-level generator that produces detailed predictions conditioned on the high-level output. This decomposition yields high perceptual fidelity while also capturing strong spatial and semantic representations. We further show that pretraining with a diffusion forcing objective yields substantially richer internal representations than the standard teacher forcing objective, while teacher forcing – predicting only the next frame from clean context – produces more stable autoregressive rollouts. We therefore introduce a generic two-stage training paradigm that pretrains the model with diffusion forcing and fine-tunes with teacher forcing, combining the representational benefits of the former with the rollout stability of the latter. Our approach achieves state-of-the-art results across the standard suite of driving world model evaluations on established benchmarks, including long-horizon generation fidelity, steering responsiveness evaluated on counterfactual scenarios, and internal representation quality. Project page with code, demo, checkpoints and qualitative results: https://lmb-freiburg.github.io/orbis2.github.io/
Feed-forward 3D reconstruction models such as DUSt3R, VGGT, and Depth Anything 3 (DA3) are transformer-based foundation models that infer camera geometry and dense scene structure in a single forward pass. Trained at scale in a supervised fashion, they raise a central question: do these models build upon geometric principles akin to traditional multi-view pipelines, or do they primarily rely on learned priors arising from the large-scale training setup? We find that epipolar geometry emerges within the intermediate layers of all three models and is causally linked to correspondence patterns in attention heads. To study this, we perform a systematic analysis of their internal representations across three real-world datasets and a controlled synthetic dataset. We quantify geometric understanding by probing intermediate features, analyzing attention patterns to identify correspondence matching patterns, and performing targeted interventions at the attention level. Further, we assess the role of learned priors by applying challenging input-level perturbations, such as occlusions, scene ambiguities, and varying camera configurations, and compare them against classical multi-stage reconstruction pipelines.
Active learning aims to reduce the high labeling cost involved in training machine learning models on large datasets by efficiently labeling only the most informative samples. Recently, deep active learning has shown success on various tasks. However, the conventional evaluation schemes are either incomplete or below par. This study critically assesses various active learning approaches, identifying key factors essential for choosing the most effective active learning method. It includes a comprehensive guide to obtain the best performance for each case, in image classification and semantic segmentation. For image classification, the AL methods improve by a large-margin when integrated with data augmentation and semi-supervised learning, but barely perform better than the random baseline. In this work, we evaluate them under more realistic settings and propose a more suitable evaluation protocol. For semantic segmentation, previous academic studies focused on diverse datasets with substantial annotation resources. In contrast, data collected in many driving scenarios is highly redundant, and most medical applications are subject to very constrained annotation budgets. The study evaluates active learning techniques under various conditions including data redundancy, the use of semi-supervised learning, and differing annotation budgets. As an outcome of our study, we provide a comprehensive usage guide to obtain the best performance for each case.
Existing world models for autonomous driving struggle with long-horizon generation and generalization to challenging scenarios. In this work, we develop a model using simple design choices, and without additional supervision or sensors, such as maps, depth, or multiple cameras. We show that our model yields state-of-the-art performance, despite having only 469M parameters and being trained on 280h of video data. It particularly stands out in difficult scenarios like turning maneuvers and urban traffic. We test whether discrete token models possibly have advantages over continuous models based on flow matching. To this end, we set up a hybrid tokenizer that is compatible with both approaches and allows for a side-by-side comparison. Our study concludes in favor of the continuous autoregressive model, which is less brittle on individual design choices and more powerful than the model built on discrete tokens. Project page with code, model checkpoints and visualization can be found here: https://lmb-freiburg.github.io/orbis.github.io/
Energy predicament is that the most vital matter in today’s world. Regular energy resources aren’t only restricted but also the prime wrongdoer for environmental pollution. Solar energy resources are becoming predominance within the world to reduce the reliance on Traditional resources. Solar energy is swiftly gaining the main target as a crucial means of inflate renewable energy operation. Solar collector those transform sun’s energy into electricity are expensive and incompetent. Different technique is applied to extend the productivity of the photovoltaic cell to scale back the value. A solar tracking system is that the most expropriate technology to reinforce the effectiveness of solar cells by trace the sun. The Proposed solar tracker has a specific control mechanism that will provide better ways of the controlling system. A small precursor of a solar tracking system is additionally assembled to implement the planning technique presented here.
This paper proposes leaky least mean fourth (LLMF) algorithm for the control of five-level series connected H-bridge multilevel inverter (5L SCHB-MLI). The SCHB multilevel inverter is a popular configuration due to its wide application for high-power and medium-voltage applications. The complete system comprising the grid, load, and compensator is modeled in MATLAB/SIMULINK. Prototype hardware is also developed in the laboratory for experimental validation. The 5L SCHB inverter is connected in a single-phase power distribution system and controlled as a compensator. The voltage across both DC links of the H-bridges is regulated to almost equal value by the proposed control algorithm. The closed loop system is designed to mitigate total harmonic distortion (THD) in the source current and improve the power quality of the proposed system. The system is analyzed under the steady state and dynamic state, and the results are also verified on the hardware prototype developed in the laboratory. The proposed algorithm is also compared with conventional least means square (LMS) algorithm and notch filter algorithm on several parameters. The paper discusses a detailed simulation and hardware analysis on the control of five-level MLI topology.
This paper presents Adaptive Radial Basis Function Neural Network (A-RBFNN) algorithm for the control of two level and five level converters. This paper investigates the performance of five level cascaded H-bridge (CHB) multilevel converter as compared to two level converters in power distribution system to improve power quality of connected system. The CHB multilevel inverter is a popular configuration because of its wide application for high power medium voltage requirement. The power system including two level (2L) and five level (5L) converter is modeled in MATLAB environment. Experimental setup is also developed in the laboratory for hardware verification purpose and both the 2L and 5L CHB converters are integrated to single phase system as compensators. The A-RBFNN algorithm is able to regulate the DC link voltages to its rated value for two H-bridges in 5L configuration. The designed algorithm is able to compensate for PQ issues of the connected system. It is also tested under steady state and dynamic conditions. Further, A-RBFNN is used to extract individual harmonics of any desired order. Finally, comparison is presented based on the performance results of the proposed algorithm with 2L as well as 5L converters.
This paper proposes an innovative approach to the design and control of Electric Vehicle Chargers (EVCs) by employing a DC-DC CUK converter. The surge in Electric Vehicle (EV) adoption necessitates efficient, reliable and grid-compatible charging infrastructure. To address this demand, a comprehensive study focused on optimizing the CUK converter for EVCs is described in this paper. This approach includes meticulous component selection, topology configuration, and advanced control algorithms to achieve superior voltage regulation and enhanced efficiency. Further, the challenges posed by input voltage variations and dynamic load fluctuations commonly encountered in EV charging scenarios is also analyzed to ensure stable and reliable operation of the EVCs. The simulation results suggest that the proposed system is feasible and practical, with higher charging efficiency and lower power losses. Additionally, this approach also enhances grid integration capabilities, aligning with the sustainability goals of modern transportation systems. In summary, review of the literature and detailed simulation on bidirectional DC-DC CUK converters has been conducted about their modes of operation. MATLAB SIMULINK software has been utilized to simulate the Steady State Analysis.
This paper presents a novel topology of Reduced Switch Five Level Inverter (RSFLI) for the integration of photovoltaic based renewable energy source and Electric Vehicle (EV) charger. The new RSFLI has simple structure with low cost due to reduced switch count and it also meets the requirement of high power, medium voltage in power plants and industries. The system design integrates grid, nonlinear load, compensator, solar panel and EV charger. The proposed system is modeled in MATLAB/SIMULINK and a Third Order Sinusoidal Integrator control algorithm is used for the control of RSFLI. The RSFLI has two separate DC links. The suggested control technique regulates both DC link voltages of the RSFLI to almost equal values. The RSFLI is utilized as a compensator and coupled to a single phase power distribution system. The complete closed loop system is designed to charge as well as discharge the EV battery to support grid. A laboratory prototype hardware has been developed for the purpose of experimental validation. The proposed system is analyzed under various dynamic conditions and the results are validated using the hardware prototype. The proposed inverter configuration is further compared with conventional inverter, cascaded H-bridge 5 level inverter on several parameters. The RSFLI along with bidirectional DC–DC converter smoothly performs G2V and V2G modes of operation. The control algorithm is tested for different dynamics such as reduction in load, change in PV output and EV. Extensive hardware study has been done to test the effectiveness of proposed system and the results are found to be satisfactory.
Significant advancements in power electronics have led to the development of a suitable platform for exploring various multilevel inverter (MLI) topologies. The paper introduces a novel asymmetrical multilevel inverter topology called “A new Criss-Cross based asymmetrically configured T-Type Multi-Level Inverter” that exhibits various beneficial features such as high-quality staircase sinusoidal output voltage, reduced number of power switches, and fewer filter requirements. The proposed topology is designed for 27 levels with a minimized number of inverter components, and its performance is evaluated using both simulation and experimental results. The simulation is conducted using MATLAB/Simulink with a sinusoidal pulse-width modulation (SPWM) technique, and the experimental results are validated using a dSPACE real-time controller. A comparative study is also conducted with other recent proposed topologies, which reveals that the proposed topology requires fewer total MLI components in terms of power switches, isolated DC sources, and main diodes. The simulation and experimental results are analyzed for two different modulation indices, i.e., 0.3 and 1. The output voltage contains 14.89 and 3.36
Active learning is particularly of interest for semantic segmentation, where annotations are costly. Previous academic studies focused on datasets that are already very diverse and where the model is trained in a supervised manner with a large annotation budget. In contrast, data collected in many driving scenarios is highly redundant, and most medical applications are subject to very constrained annotation budgets. This work investigates the various types of existing active learning methods for semantic segmentation under diverse conditions across three dimensions - data distribution w.r.t. different redundancy levels, integration of semi-supervised learning, and different labeling budgets. We find that these three underlying factors are decisive for the selection of the best active learning approach. As an outcome of our study, we provide a comprehensive usage guide to obtain the best performance for each case. We also propose an exemplary evaluation task for driving scenarios, where data has high redundancy, to showcase the practical implications of our research findings.
Water resources are crucial for various human needs, including preventive health, food production, energy generation, ecological balance, and socioeconomic development. However, a significant portion of the global population, particularly in water-poor regions, lacks access to clean and sufficient drinking water. This paper presents a sustainable solution implemented with photovoltaic water pump system to address these challenges. This paper reviews relevant literature on photovoltaic technology, maximum power point tracking, and solar cell technologies. It also outlines the basic of BLDC motor for these applications including the functioning of the proposed photovoltaic water pump system. This conference paper explores the use of DCDC Boost Converter in photovoltaic (PV) systems to achieve Maximum Power Point Tracking (MPPT). The Boost Converter is discussed in detail, highlighting its significance in voltage boosting and power optimization for PV panels. Impedance matching is emphasized as a unique advantage of DC-DC Boost Converters, enabling the extraction of maximum power from PV arrays. The paper also presents the modes of operation for the Boost Converter, focusing on charging and discharging modes.
The Coordinating directional overcurrent relays (DOCRs) in a mesh system with various sources is a difficult and time-consuming restricted optimization problem. Over the last decade, DOCR coordination has been primarily done manually. However, recent initiatives have attempted to overcome this by treating the coordination of directed overcurrent relays in an electric power system as an optimization problem. The major purpose is to identify the best configuration for DOCRs, which minimizes relay operating time for faults within their assigned protection zones while ensuring proper relay coordination. The primary goal of this coordination challenge is to avoid relay breakdowns and unnecessary separation of the system’s healthy components. To address this issue, a methodology has been created that entails constructing an algorithm capable of coordinating directional overcurrent relays using an optimization approach. This systematic method attempts to improve the efficiency and accuracy of DOCR coordination by streamlining the process and decreasing the human work that has historically been associated with it.
Active learning automatically selects samples for annotation from a data pool to achieve maximum performance with minimum annotation cost. This is particularly critical for semantic segmentation, where annotations are costly. In this work, we show in the context of semantic segmentation that the data distribution is decisive for the performance of the various active learning objectives proposed in the literature. Particularly, redundancy in the data, as it ap-pears in most driving scenarios and video datasets, plays a large role. We demonstrate that the integration of semi-supervised learning with active learning can improve performance when the two objectives are aligned. Our experimental study shows that current active learning benchmarks for segmentation in driving scenarios are not realistic since they operate on data that is already curated for maximum diversity. Accordingly, we propose a more realistic evaluation scheme in which the value of active learning becomes clearly visible, both by itself and in combination with semi-supervised learning.
Intelligent sampling from simulation becomes crucial due to storage and hardware constraints. This research focuses on developing an intelligent acquisition strategy for synthetic data and evaluates multiple approaches to address the limitations of existing domain adaptation methods. Selecting suitable synthetic data for real-world model training presents challenges, as accurately representing the real world remains elusive. We tackle the task of adapting from synthetic to real-world data through unsupervised domain adaptation, a challenging setting for perception systems. The performance of our acquisition function is measured by its facilitation of this adaptation.We showcase different strategies, to assign value to synthetic images. Acquisition functions either operate based on synthetic data alone or take the given real world target domain into account, to assign a value to synthetic images. Leveraging assumptions from semi-supervised learning, we identify challenging real-world images and find their counterparts in the synthetic world. Evaluation is conducted using the GTA-5 dataset as the representative synthetic world and the Cityscapes and ACDC dataset as the target do-main. State-of-the-art unsupervised domain adaptation approaches are employed to assess the effectiveness of our acquisition function.By advancing the utilization of synthetic data in training perception systems, this research contributes to improved real-world performance. Our findings demonstrate the potential of intelligent acquisition strategies for enhancing the adaptation from synthetic to real-world domains.
Vision-language modeling has enabled open-vocabulary tasks where predictions can be queried using any text prompt in a zero-shot manner. Existing open-vocabulary tasks focus on object classes, whereas research on object attributes is limited due to the lack of a reliable attribute-focused evaluation benchmark. This paper introduces the Open-Vocabulary Attribute Detection (OVAD) task and the corresponding OVAD benchmark. The objective of the novel task and benchmark is to probe object-level attribute information learned by vision-language models. To this end, we created a clean and densely annotated test set covering 117 attribute classes on the 80 object classes of MS COCO. It includes positive and negative annotations, which enables open-vocabulary evaluation. Overall, the benchmark consists of 1.4 million annotations. For reference, we provide a first baseline method for open-vocabulary attribute detection. Moreover, we demonstrate the benchmark's value by studying the attribute detection performance of several foundation models.
This paper proposes a new Normalised Adaptive Regression Least Mean Mixed Norm (NARLMMN) algorithm to control three phase Voltage Source Converter (VSC). Further, the paper considers an On-Board EV Charger and a PV source connected at DC link of the VSC. The PV feeds real power to Electric Vehicle (EV) and the grid. Both the Grid to Vehicle (G2V) and Vehicle-to-Grid (V2G) operations are implemented using bidirectional Buck-Boost converter which is an emergent area of research. The grid supplies power to charge the battery of EV in G2V mode of operation. In V2G mode of operation, the stored energy in the EV battery bank is reutilized to supply back to the utility grid, which helps in peak shaving, load balancing, voltage management and improving the system reliability. The bidirectional DC-DC converter, along with the battery bank are connected at the DC link of three phase VSC and controlled to implement charging and discharging modes. The performance of PV with VSC and PV, EV with VSC is analysed on MATLAB/Simulink platform. The PWM control technology is used to regulate the voltage and current used for charging and discharging batteries. The dynamics of the proposed system are analyzed and results are shown in this paper with the new proposed algorithm.
Electric vehicle (EV) technology is developing at a very fast pace. The Vehicle to grid (V2G) and Grid to vehicle (G2V) technology enables bidirectional power transfer between an electric vehicle and the grid. It is an emergent area of research. In the G2V operation mode EV, batteries are charged from the grid. The energy stored in the batteries may also be supplied back to the power grid during the V2G operation mode, which helps in maximum demand saving, load balancing voltage management and improved system reliability. An onboard bilateral battery charger for EV is proposed in this research paper, with G2V and V2G applications. Bi-directional power electronics-based converter is interfaced between EV and the grid to provide G2V and V2G modes. A Second Order Generalized Integrator (SOGI) control technique is used to control the H-bridge inverter which shows stable steady state and good dynamic performance. The battery charging/discharging current and voltage are controlled using a PWM controller further effect of nonlinear load dynamics on the system performance is also studied. Exhaustive simulation study is performed in MATLAB/Simulink environment which is also reflected in this paper.