Flow Matching (FM) has recently emerged as a powerful approach for high-quality visual generation. However, their prohibitively slow inference due to a large number of denoising steps limits their potential use in real-time or interactive applications. Existing acceleration methods, like distillation, truncation, or consistency training, either degrade quality, incur costly retraining, or lack generalization. We propose FlowCast, a training-free speculative generation framework that accelerates inference by exploiting the fact that FM models are trained to preserve constant velocity. FlowCast speculates future velocity by extrapolating current velocity without incurring additional cost, and accepts it if it is within a mean-squared error threshold. This constant-velocity forecasting allows redundant steps in stable regions to be aggressively skipped while retaining precision in complex ones. FlowCast is a plug-and-play framework that integrates seamlessly with any FM model and requires no auxiliary networks. We also present a theoretical analysis and bound the worst-case deviation between speculative and full FM trajectories. Empirical evaluations demonstrate that FlowCast achieves $>2.5\times$ speedup in image generation, video generation, and editing tasks, outperforming existing baselines with no quality loss as compared to standard full generation.
Retrieval-Augmented Generation (RAG) is often used with Large Language Models (LLMs) to infuse domain knowledge or user-specific information. In RAG, given a user query, a retriever extracts chunks of relevant text from a knowledge base. These chunks are sent to an LLM as part of the input prompt. Typically, any given chunk is repeatedly retrieved across user questions. However, currently, for every question, attention layers in LLMs fully compute the Keys and Values (KVs) repeatedly for the input chunks, as state-of-the-art methods cannot reuse KV-caches when chunks appear at arbitrary locations or with arbitrary contexts. Naive reuse leads to output quality degradation. This leads to potentially redundant computations on expensive GPUs and increases latency. In this work, we propose Cache-Craft, a system for managing and reusing precomputed KVs corresponding to the text chunks (which we call chunk-caches) in RAG-based systems. We present how to identify chunk-caches that are reusable, how to efficiently perform a small fraction of recomputation to fix the cache and maintain output quality, and how to efficiently store and evict chunk-caches in the hardware for maximizing reuse while masking any overheads. With real production workloads as well as synthetic datasets, we show that Cache-Craft reduces redundant computation by 51% over SOTA prefix-caching and 75% over full recomputation. Additionally, with continuous batching on a real production workload, we get a 1.6× speed up in throughput for both the LLama-3-8B and 70B models and a 2.1× and 2× reduction in end-to-end response latency respectively, compared to prefix-caching, while maintaining generation quality.
Personalized preference alignment for LLMs with diverse human preferences requires evaluation and alignment methods that capture pluralism. Most existing preference alignment datasets are logged under policies that differ substantially from the evaluated LLMs, and existing off-policy estimators focus solely on overall utility while ignoring preference pluralism. Extending Off-Policy Evaluation (OPE) to pluralistic preference alignment, therefore, remains an open question. Thus, we propose the Pluralistic Off-Policy Evaluation (POPE), the first framework for offline pluralistic preference evaluation and alignment in LLMs. POPE includes a unified reward function that combines (1) a collaborative utility component derived from human preference signals (e.g., upvotes or relevance scores) and (2) a diversity component inspired by entropy-based coverage measures, together reflecting pluralistic alignment. Furthermore, to estimate this reward from logged interactions, we derive decomposable inverse propensity scoring (IPS) estimators that separately evaluate relevance and diversity. Theoretically, we prove that our decomposed IPS estimators establish a lower bound on their variance. With the off-policy evaluated value function, we can directly enable off-policy optimization to further enhance pluralistic alignment. Empirical results demonstrate that POPE efficiently enhances pluralistic response generation and maintains the models' general capabilities on downstream tasks
Traditional ML models utilize controlled approximations during high loads, employing faster, but less accurate models in a process called accuracy scaling. However, this method is less effective for generative text-to-image models due to their sensitivity to input prompts and performance degradation caused by large model loading overheads. This work introduces a novel text-to-image inference system that optimally matches prompts across multiple instances of the same model operating at various approximation levels to deliver high-quality images under high loads and fixed budgets.
The purpose of the study is to highlight the important amenities listing variables and the pervasiveness of the determinants to the "superhost" status in the online shared destination accommodation of the tourism and hospitality market. The potential guests consider the superhost status as a quality indicator of the Airbnb accommodation and promise of the host for their offered amenities and thus increasing their rental demands resulting in more revenue from more bookings. Using the Airbnb listing dataset of four cities of Canada and eleven cities of the United States, the study applied six different data mining techniques to find the importance of listing variables. It identified essential amenities based on their presence among the top variables in the applied models. The findings of the study were the identification of a reduced set of offered amenities and service variables that influence Airbnb's reward of the "superhost" badge as a visual symbol of trust and credibility. Thus, the findings help to develop better service quality for tourists and at the same time helps to mitigate the complexities confronted by hosts with multiple listings.
Text-to-image diffusion models excel in generating photo-realistic images but are hampered by slow processing times. Training-free retrieval-based acceleration methods, which leverage pre-generated "trajectories," have been introduced to address this. Yet, these methods often lack diversity and fidelity as they depend heavily on similarities to stored prompts. To address this, we present ReCON (Retrieving CONcepts), an innovative retrieval-based diffusion acceleration method that extracts visual "concepts" from prompts, forming a knowledge base that facilitates the creation of adaptable trajectories. Consequently, ReCON surpasses existing retrieval-based methods, producing high-fidelity images and reducing required Neural Function Evaluations (NFEs) by up to 40%. Extensive testing on MS-COCO, Pick-a-pick, and DiffusionDB datasets confirms that RECON consistently outperforms established methods across multiple metrics such as Pick Score, CLIP Score, and Aesthetics Score. A user study further indicates that 76% of images generated by ReCON are rated as the highest fidelity, outperforming two competing methods, a purely text-based retrieval and a noise similarity-based retrieval.
Text-to-image generation using diffusion models has seen explosive popularity owing to their ability in producing high quality images adhering to text prompts. However, production-grade diffusion model serving is a resource intensive task that not only require high-end GPUs which are expensive but also incurs considerable latency. In this paper, we introduce a technique called approximate-caching that can reduce such iterative denoising steps for an image generation based on a prompt by reusing intermediate noise states created during a prior image generation for similar prompts. Based on this idea, we present an end to end text-to-image system, Nirvana, that uses the approximate-caching with a novel cache management-policy Least Computationally Beneficial and Frequently Used (LCBFU) to provide % GPU compute savings, 19.8% end-to-end latency reduction and 19% dollar savings, on average, on two real production workloads. We further present an extensive characterization of real production text-to-image prompts from the perspective of caching, popularity and reuse of intermediate states in a large production environment.
In this paper, we investigate the effect of organizational values/culture on sustainable practices in Indian SMEs, and observe that organizational values/culture positively affect waste disposal/recycling, and employee-related social practices. Further, employee-related social practices act as a mediating variable between organizational values/culture, and firms' environmental and community-related social practices. We also examine the moderating role of family influence and observe that for family SMEs, the effects of organizational values/culture on waste disposal/recycling, and employee-related social practices are stronger than those for non-family SMEs. The paper concludes by highlighting the implications and limitations of the study, and possible directions for future research.
A 335 sq km triangle, marked by the positions of Limpiyadhura-Kalapani-Lipu Lekh, currently in India's possession and claimed by Nepal, has developed into an open feud between the two countries. Although the dispute has been festering, it is the timing of the recent flare-up, its connection with other border tensions and the unprecedented actions taken by the Nepalese government that merit attention. Given Nepal's physical location, the crisis inevitably has a China factor especially as China's economic influence in the country has been growing rapidly. In 2020, pre-Covid 19 figures indicated that China accounted for over 90% of foreign direct investments. To overcome Nepal's dependence on India not only for imports but also on Indian ports of entry, Kathmandu signed a transit protocol with Beijing last year. The current flare up between India and China along the un-demarcated border known as the Line of Actual Control (LAC), leading to the death of 20 Indian soldiers and an unspecified number of their Chinese adversaries – the first incident involving death of troops at this scale – has given a special piquancy to the border issue between India and Nepal.On May 20th 2020 Nepal published a new official map including the above region as part of its territory, a decision that was pitched as being a direct response to Indian government actions. On May 8th India had inaugurated a newly-built road link to Kaishal Mansarovar in the Tibetan Autonomous Region that runs through the Lipu lekh pass. Even earlier, in November 2019 the Nepalese government had objected to India's 'new' political map, released after the internal re-organisation of boundaries and status of Jammu and Kashmir. Nepal had then protested the inclusion of Kalapani, a 35 square kilometre area in the Pithoragarh district under the control of the Indo-Tibetan Border Police, as part of India. However, it must be pointed out that maps since 1905, released by the Survey of India, the national survey and mapping organisation, have shown this area as Indian territory. Furthermore, India has had an army base in the Kalapani region near the Lipu lekh pass since the early 1950s and despite requests, has not given up this position due to its strategic value. The high ground at the pass enables the Indian army to monitor routes that connect with Tibet.Nepal's latest moves to assert its position and to directly challenge its neighbour have included a decision by the governing party to table a bill in Parliament to amend the constitution and update the new political map as part of the national emblem. The lower and upper houses have both unanimously endorsed this proposal. Describing Nepal's new official map as "artificial" and unacceptable, Indian government officials have portrayed the actions as unilateral. What explains this rapid deterioration in relations and the willingness of the Nepalese regime to escalate tensions with its neighbour and to challenge the status quo? While it is true that India had chosen to ignore the problem, the current hard-line stance on the part of the Nepalese government has reduced the room to manoeuvre in finding a resolution through dialogue.
This paper explores the role of crude oil in determining corn prices for data on the weekly front future prices in the United States. With 38% of corn production allocated toward fuel ethanol, a possible effect of crude oil price variation on corn price fluctuations is theoretically indicated. To test this theory, two complementary approaches—a parametric multiple regression and a non-parametric multivariate adaptive regression splines approach are employed. Along with indicating a weak relationship between corn and crude oil prices, the results suggest that corn price responds nonlinearly to the changes in soybean and wheat prices.
Automated visualization recommendation (Vis-Rec) models help users to derive crucial insights from new datasets. Typically, such automated Vis-Rec models first calculate a large number of statistics from the datasets and then use machine-learning models to score or classify multiple visualizations choices to recommend the most effective ones, as per the statistics. However, state-of-the-art models rely on a very large number of expensive statistics and therefore using such models on large datasets becomes infeasible due to prohibitively large computational time, limiting the effectiveness of such techniques to most large real-world datasets. In this paper, we propose a novel reinforcement-learning (RL) based framework that takes a given Vis-Rec model and a time budget from the user and identifies the best set of input statistics, specifically for a target dataset, that would be most effective while generating accurate enough visual insights. We show the effectiveness of our technique as it enables two state of the art Vis-Rec models to achieve up to 10X speedup in time-to-visualize on four large real-world datasets.
This paper aims to examine the determinants of life insurance consumption in 30 OECD countries using panel data from 1996 to 2020. This study uses GDP per capita, Life expectancy, Urbanization, School education, and Health expenditure as the determinants to measure the OECD countries’ life insurance consumption. Insurance density is used as a proxy for life insurance consumption. Fully Modified Ordinary Least Squares (FMOLS), Dynamic Ordinary Least Squares (DOLS), and causality tests are applied in this study. Our empirical results revealed that the variables urbanization, school education, and GDP per capita significantly impact life insurance consumption, whereas life expectancy and health expenditure were found to have an insignificant relationship in estimating life insurance consumption. These findings will help all insurance industry stakeholders in OECD countries in policy formulation and decision making.
Although the used car market in India is enormous, with an annual 27.1 billion USD worth of car sales, no academic study examining the pricing of Indian used cars is available. The average on-road life of cars in India is high, exceeding twenty years and pass through multiple hands. The price of old cars is also a primary determinant of the prices of new cars. The article discusses the non-linear reduction of used car prices on account of age, kilometer driven, and the number of previous owners of the cars. In addition to using the Ordinary Linear Squares (OLS), we also applied Multivariate Adaptive Regression Splines (MARS) techniques to capture non-linear price relationships with other depreciating variables. The resulting error estimation from these two methods demonstrates that the complexity of non-linear modeling using MARS significantly improves fitment accuracy.
In this paper, we study personalized federated learning for text classification with Pretrained Language Models (PLMs). We identify two challenges in efficiently leveraging PLMs for personalized federated learning: 1) Communication. PLMs are usually large in size, inducing huge communication cost in a federated setting. 2) Local Training. Training with PLMs generally requires back-propagation, during which memory consumption can be several times that of the forward-propagation. This may not be affordable when the PLMs are trained locally on the clients that are resource constrained, e.g., mobile devices with limited access to memory resources. In solving these, we propose a training framework that includes an approach of discrete local search for gradient-free local training, along with a compression mechanism inspired from the linear word analogy that allows communicating with discretely indexed tokens, thus significantly reducing the communication cost. Experiments show that our gradient-free framework achieves superior performance compared with baselines.
Predicting profitable customers is a strategic knowledge portfolio of retailer managers because some customers are better profitable than others in a business. The present work is an effort to demonstrate a better model of predicting profitable customers. We apply the k-means algorithm to identify customer patterns based on Recency, Frequency, and Monetary (RFM) attributes computed from a real-life dataset of UK-based and registered non-store online retail. Six data mining models have been applied to each identified pattern and overall data to predict whether each customer would purchase in the next six months or not. A comparative analysis of identified pattern characteristics and predictable performances and Type I and Type II errors have been performed to identify the target customer group in terms of better predictability and profitability. The identified patterns help to generate novel marketing strategies. Thus, the retailers may successfully target the most consistently profitable customer groups to apply diverse knowledge on marketing strategies for the specific pattern.
The analysis is aimed at comprehending the interplay between exchange rates, hotel pricing, occupancy, and revenue per available room in three Nordic countries (Finland, Norway, and Sweden) using a novel approach of dynamic common correlated effects that can handle cross-sectional interdependencies, variations, and dynamics within the data. The paper finds that exchange rate appreciation has a significant adverse outcome on hotel occupancy rates but a positive effect on revenue per available room. In contrast, exchange rate fluctuations did not significantly impact hotel room prices, implying that hotels in Nordic countries can effectively consider the asymmetric influences of exchange rate fluctuations in setting room prices. The paper discusses implications for hotel managers and policymakers in coping with exchange rate fluctuations and optimizing hotel pricing.
State-of-the-art video database management systems (VDBMSs) often use lightweight proxy models to accelerate object retrieval and aggregate queries. The key assumption underlying these systems is that the proxy model is an order of magnitude faster than the heavyweight oracle model. However, recent advances in computer vision have invalidated this assumption. Inference time of recently proposed oracle models is on par with or even lower than the proxy models used in state-of-the-art (SoTA) VDBMSs. This paper presents Seiden, a VDBMS that leverages this radical shift in the runtime gap between the oracle and proxy models. Instead of relying on a proxy model, Seiden directly applies the oracle model over a subset of frames to build a query-agnostic index, and samples additional frames to answer the query using an explorationexploitation scheme during query processing. By leveraging the temporal continuity of the video and the output of the oracle model on the sampled frames, Seiden delivers faster query processing and better query accuracy than SoTA VDBMSs. Our empirical evaluation shows that Seiden is on average 6.6 x faster than SoTA VDBMSs across diverse queries and datasets.
The welfare state, once seen as the best institutional response to people in need, has steadily come under pressure, as much from shrinking state capacities as from neo-liberal advocates of individual responsibility. Still, despite decline of the post-war consensus on the efficacy of the welfare state, social ‘vulnerability’ still remains the key focus of public policy. However, though much in use in contemporary political discourse, the logical and practical implications of social vulnerability remain unclear. Its essential subjectivity – it is the ‘feeling of vulnerability’ which makes one vulnerable – turns the concept into a catch-all variable, impeding rigorous theoretical and empirical analysis. I respond to this problem with a ‘vulnerability-responsibility’, model. Its parameters include a ‘responsive’ state, an active civil society and a participatory political environment, bolstered by the assertion of agency of the vulnerable. With India as an empirical exemplar, the essay shows how ‘nested’ vulnerability – a community of pro-active citizens in need of urgent and vital assistance - in the backdrop of a responsive state and competitive, robust and resilient political participation, can generate a sustainable, context-relevant process to cope with the problem of social vulnerability. The model, currently aimed at vulnerable citizens in a democratic state, has the potential of being extended to non-democracies as well as vulnerable non-citizens into its domain.