Due to adverse factors such as varying illumination, noise, and imaging artifacts, achieving finegrained image segmentation of objects remains a significant challenge. To address this, we propose a level set method based on global alternating minimization. Specifically, a total variation (TV) regularization term weighted by a gradient-based edge indicator function is incorporated into a convex energy functional, enhancing the model's ability to detect weak edges. Subsequently, an efficient segmentation framework is constructed based on the Alternating Direction Method of Multipliers (ADMM), providing a closed-form solution that improves both numerical stability and convergence speed. By adopting a convex optimization scheme, the proposed model eliminates explicit time-step dependence, thereby improving adaptability and flexibility in the temporal domain. Experimental results demonstrate that the proposed method possesses a global minimization property and consistently outperforms state-of-the-art segmentation models on publicly available datasets. Notably, compared to the Segment Anything Model (SAM), the proposed method reduces the maximum CT measurement error of the ball-plate standard by 65.66 %.
In this paper, we present the first constant-approximation algorithm for budgeted sweep coverage problem (BSC). The BSC involves designing routes for a number of mobile sensors (a.k.a. robots) to periodically collect information as much as possible from points of interest (PoIs). To approach this problem, we propose to first examine the multi-orienteering problem (MOP). The MOP aims to find a set of m vertex-disjoint paths that cover as many vertices as possible while adhering to a budget constraint B. We develop a constant-approximation algorithm for MOP and utilize it to achieve a constant-approximation for BSC. Our findings open new possibilities for optimizing mobile sensor deployments and related combinatorial optimization tasks.
Digital platforms increasingly rely on personalized ranking algorithms to drive engagement and monetization, where product displays are tailored simultaneously by who the customer is and what the products offer. A central operational challenge lies in learning these customer--product interaction patterns online from random and partial feedback, as customer attention is inherently limited. Motivated by the widely observed cascade browsing behavior in media platforms and online marketplaces, where customers scan ranked recommendations top-down and click the first appealing one, we study the dynamic product ranking problem under dual contextual information: customer characteristics and product attributes. Existing dynamic contextual ranking approaches typically treat customer and product contexts as separate feature vectors, which neglects the interaction structure between them. In many applications, however, customer preferences for products are driven by a small number of latent factors, implying a low-dimensional structure between customer--product interactions. We capture this structure by embedding the dual contexts into a unified feature matrix and imposing a low-rank structure on the parameter matrix that governs customer choice behavior. Building on this representation, we formulate a generalized low-rank cascading bandit framework that leverages both sequential customer feedback and the structured interaction to enable effective online learning. Within this framework, we first consider a base setting focused on click rate maximization that a common objective as in movie streaming services and content platforms, and establish optimal regret bounds under mild conditions. We then explore an extension in which products generate heterogeneous revenues per click, as in sponsored search and advertising platforms. This gives rise to a revenue maximization problem with a non-trivial combinatorial ranking structure at each round. To address this challenge, we develop a backward dynamic programming approach integrated into our framework for joint learning and optimization. Numerical experiments on both simulated and real-world datasets demonstrate that the proposed algorithms consistently outperform strong benchmarks across both click rate and revenue maximization tasks.
Federated Prototype Learning (FedPL) has emerged as an effective strategy for handling data heterogeneity in Federated Learning (FL). In FedPL, clients collaboratively construct a set of global feature centers (prototypes), and let local features align with these prototypes to mitigate the effects of data heterogeneity. The performance of FedPL highly depends on the quality of prototypes. Existing methods assume that larger inter-class distances among prototypes yield better performance, and thus design different methods to increase these distances. However, we observe that while these methods increase prototype distances to enhance class discrimination, they inevitably disrupt essential semantic relationships among classes, which are crucial for model generalization. This raises an important question: how to construct prototypes that inherently preserve semantic relationships among classes? Directly learning these relationships from limited and heterogeneous client data can be problematic in FL. Recently, the success of pre-trained language models (PLMs) demonstrates their ability to capture semantic relationships from vast textual corpora. Motivated by this, we propose FedTSP, a novel method that leverages PLMs to construct semantically enriched prototypes from the textual modality, enabling more effective collaboration in heterogeneous data settings. We first use a large language model (LLM) to generate fine-grained textual descriptions for each class, which are then processed by a PLM on the server to form textual prototypes. To address the modality gap between client image models and the PLM, we introduce trainable prompts, allowing prototypes to adapt better to client tasks. Extensive experiments demonstrate that FedTSP mitigates data heterogeneity while significantly accelerating convergence.
Users interacting with large language models (LLMs) under their real identifiers often unknowingly risk disclosing private information. Automatically notifying users whether their queries leak privacy and which phrases leak what private information has therefore become a practical need. Existing privacy detection methods, however, were designed for different objectives and application domains, typically tagging personally identifiable information (PII) in anonymous content, which is insufficient in real-name interaction scenarios with LLMs. In this work, to support the development and evaluation of privacy detection models for LLM interactions that are deployable on local user devices, we construct a large-scale multilingual dataset with 249K user queries and 154K annotated privacy phrases. In particular, we build an automated privacy annotation pipeline with strong LLMs to automatically extract privacy phrases from dialogue datasets and annotate leaked information. We also design evaluation metrics at the levels of privacy leakage, extracted privacy phrase, and privacy information. We further establish baseline methods using light-weight LLMs with both tuning-free and tuning-based methods, and report a comprehensive evaluation of their performance. Evaluation results demonstrate that the 1B model, fine-tuned with our dataset, outperforms the directly prompted 72B model. However, a gap remains between current performance and the requirements of real-world LLM applications, motivating future research into more effective local privacy detection methods grounded in our dataset(1).
Heterogeneous Graph Neural Networks excel in various recommendation scenarios by effectively modeling and leveraging diverse information. However, two key challenges remain. First, heterogeneous information, such as user-item interactions, social relationships, and category tags, is found in distinct semantic spaces. Directly merging this information can lead to semantic confusion, making it difficult for the model to differentiate between relationships and reducing recommendation accuracy. Second, each type of heterogeneous relationship contains unique semantic characteristics. Current methods often focus solely on connectivity, neglecting these unique semantics, which limits the model’s ability to understand and represent heterogeneous information effectively. To address these challenges, we propose a novel approach named Multi-view Heterogeneous Graph with Cross-view Projection (MHGCP). This approach creates independent views for each heterogeneous semantic type to mitigate semantic confusion. Additionally, it introduces a cross-view projection layer that facilitates information transfer between semantic views and encodes inter-view relationships, allowing the model to indirectly capture the unique properties of each view. We tested our model on three real datasets, demonstrating superior performance. Through ablation studies and case studies, we validated the contribution of key modules in our approach to performance improvement. The implementation of the model can be found on https://github.com/fxl951677676/MHGCP .
Large language models (LLMs) offer strong capabilities but raise cost and privacy concerns, whereas small language models (SLMs) facilitate efficient and private local inference yet suffer from limited capacity. To synergize the complementary strengths, we introduce a dynamic collaboration framework, where an SLM learns to proactively decide how to request an LLM during multi-step reasoning, while the LLM provides adaptive feedback instead of acting as a passive tool. We further systematically investigate how collaboration strategies are shaped by SLM and LLM capabilities as well as efficiency and privacy constraints. Evaluation results reveal a distinct scaling effect: stronger SLMs become more self-reliant, while stronger LLMs enable fewer and more informative interactions. In addition, the learned dynamic collaboration strategies significantly outperform static pipelines and standalone inference, and transfer robustly to unseen LLMs.
Clinical fusion of Single Photon Emission Computed Tomography Myocardial Perfusion Imaging (SPECT MPI) and Computed Tomography Angiography (CTA) remains limited by cross-modality misregistration and reliance on manual landmarks, which can hinder accurate ischemia localization and lesion-level functional assessment. To address this issue, we propose a registration and fusion framework for SPECT MPI and CTA that integrates functional and structural information for comprehensive cardiac evaluation. The proposed pipeline performs U-Net-based segmentation on both modalities. On SPECT MPI, only the left ventricle (LV) is extracted, and anatomical landmarks are automatically derived from characteristic LV structures. On CTA, both ventricles are segmented, and their spatial relationship is used to automatically define landmarks at the interventricular septal junction. Scale-space consistency preprocessing and landmark-driven coarse registration are applied to mitigate initial misalignment. Based on this initialization, multiple fine registration methods are evaluated on LV epicardial surface point clouds, including ICP, SICP, CPD, CluReg, FFD, and BCPD-plus-plus. The resulting transformations are then propagated to voxel-level resampling for high-precision SPECT-CTA fusion. In a retrospective cohort of 60 patients, the proposed framework preserved sub-millimeter coronary detail from CTA while accurately overlaying quantitative SPECT perfusion. Among the evaluated methods, BCPD-plus-plus achieved the highest accuracy with a mean point cloud distance of 1.7 mm. By combining robust initialization, comparative fine registration, and voxel-level fusion, the proposed approach provides a practical solution for myocardial ischemia localization and functional evaluation of coronary lesions, while remaining independent of any specific fine registration algorithm.
Embedding deep neural networks (NNs) into mixed-integer programs (MIPs) is attractive for decision making with learned constraints, yet state-of-the-art monolithic linearisations blow up in size and quickly become intractable. In this paper, we introduce a novel dual-decomposition framework that relaxes the single coupling equality u=x with an augmented Lagrange multiplier and splits the problem into a vanilla MIP and a constrained NN block. Each part is tackled by the solver that suits it best-branch and cut for the MIP subproblem, first-order optimisation for the NN subproblem, so the model remains modular, the number of integer variables never grows with network depth, and the per-iteration cost scales only linearly with the NN size. On the public SurrogateLIB benchmark, our method proves scalable, modular, and adaptable: it runs 120x faster than an exact Big-M formulation on the largest test case; the NN sub-solver can be swapped from a log-barrier interior step to a projected-gradient routine with no code changes; and swapping the MLP for an LSTM backbone still completes the full optimisation in 47s without any bespoke adaptation.
In classical auction theory with an exogenously fixed number of bidders, auctioneers maximize revenue by running a second-price auction with an appropriately chosen reserve price. Variants of this auction format are widely adopted by online auction platforms. However, empirical evidence shows that when buyer entry is endogenous, sellers tend to refrain from setting high starting prices to preserve participation. Recently, guarantee-price auctions that maintain both sides' participation have been successfully adopted by a large online auction platform, in which the platform commits to a minimum payment to the seller regardless of the auction outcome. However, the economic logic underlying this new auction format remains unclear. In this paper, we characterize the platform's optimal guarantee-pricing strategy, identify the conditions under which such a mechanism is profitable, and clarify the trade-offs involved in its adoption. We find that when the seller's entry cost is relatively low, starting-price auctions are more profitable for the platform. For intermediate entry costs, the seller does not enter under starting-price auctions, and the platform prefers the guarantee-price auction. When the seller's entry cost is sufficiently high, the platform cannot earn a profit under either auction format. From the users' perspective, buyers always weakly prefer guarantee-price auctions, while sellers weakly prefer starting-price auctions. Furthermore, we study a hybrid mechanism in which the platform both allows sellers to set a starting price and offers a guarantee. We show that the hybrid mechanism Pareto dominates either single-instrument format, as it allows the platform to flexibly deploy the most effective instrument across different market conditions. Our results provide actionable managerial insights for platform operators on when and why guaranteed price mechanisms should be implemented, with implications for both established and emerging marketplaces.
Multi-turn dialogue is the predominant form of interaction with large language models (LLMs). While LLM routing is effective in single-turn settings, existing methods fail to maximize cumulative performance in multi-turn dialogue due to interaction dynamics and delayed rewards. To address this challenge, we move from myopic, single-turn selection to long-horizon sequential routing for multi-turn dialogue. Accordingly, we propose DialRouter, which first performs MCTS to explore dialogue branches induced by different LLM selections and collect trajectories with high cumulative rewards. DialRouter then learns a lightweight routing policy from search-derived data, augmented with retrieval-based future state approximation, enabling multi-turn routing without online search. Experiments on both open-domain and domain-specific dialogue tasks across diverse candidate sets of both open-source and closed-source LLMs demonstrate that DialRouter significantly outperforms single LLMs and existing routing baselines in task success rate, while achieving a superior performance-cost trade-off when combined with a cost-aware reward.
Heterogeneous federated learning (HtFL) aims to enable collaboration among clients that differ in both data distributions and model architectures. Prototype-based methods, which communicate class-level feature centers (prototypes) instead of full model parameters, have recently shown strong potential for HtFL. Existing prototype-based HtFL methods typically reuse the MSE-based or cosine-based alignment mechanism developed for homogeneous FL when aligning client-specific representations with global prototypes. These approaches are essentially coordinate alignment, where representations of clients are forced to match the global prototypes in the embedding space in an element-wise manner. Such alignment implicitly assumes that all clients should map their representations into the feature subspace defined by the global prototypes. This assumption is reasonable in homogeneous FL, where all clients share the same feature extractor. However, it becomes problematic in HtFL, since heterogeneous feature extractors naturally induce client-specific feature subspaces, and forcing all clients to optimize within a single global subspace unnecessarily suppresses their learning capacity. We observe that coordinate alignment implicitly couples two distinct objectives: aligning inter-class semantic structure, which is directly beneficial for classification, and enforcing a shared feature basis, which is unnecessary and even harmful under model heterogeneity. Building on this insight, we design FedSAF, which shifts the alignment objective from absolute coordinates to inter-class relational structure. We demonstrate that structural alignment consistently outperforms coordinate alignment in heterogeneous settings. Experiments on multiple benchmarks show that our structural alignment outperforms state-of-the-art prototype-based HtFL methods by up to 3.52\%.
Despite the assortment optimization problem has been widely studied in the past decades, the interplay between advertising and its implications for this issue remains under-explored. This study seeks to bridge this research gap by tackling the combined challenge of advertising and assortment optimization. We assume that advertising can increase the awareness of specific products, and the magnitude of this effect is jointly depends on the product-specific effectiveness of advertising and the allocated advertising budget. For this joint problem, our objective is to maximize the expected revenue by finding the optimal advertising strategy and the displayed assortment. In this work, we analyze the structure of this problem and propose efficient approaches to solve it across different scenarios. In the unconstrained setting, we demonstrate that the optimal assortment includes products whose revenue exceeds a certain threshold. When there is a cardinality constraint for the assortment, we consider a relaxed problem and propose an efficient method to identify a near-optimal solution. We also examine the joint assortment, pricing, and advertising problem in both unconstrained and cardinality-constrained settings, incorporating the fairness constraint for the advertising strategy and extending our findings to account for consumer sequential decision-making patterns. Through a series of numerical tests, we confirm the validity of our methods and demonstrate that they outperform existing heuristic approaches.
We consider the noisy matrix sensing problem in the over-parameterization setting, where the estimated rank r is larger than the true rank r_⋆ of the target matrix X_⋆. Specifically, our main objective is to recover a matrix X_⋆∈ℝ^n_1 × n_2 with rank r_⋆ from noisy measurements using an over-parameterized factorization LR^⊤, where L ∈ℝ^n_1 × r, R ∈ℝ^n_2 × r and min{n_1, n_2}≥ r > r_⋆, with r_⋆ being unknown. Recently, preconditioning methods have been proposed to accelerate the convergence of matrix sensing problem compared to vanilla gradient descent, incorporating preconditioning terms (L^⊤ L + λ I)^-1 and (R^⊤ R + λ I)^-1 into the original gradient. However, these methods require careful tuning of the damping parameter λ and are sensitive to step size. To address these limitations, we propose the alternating preconditioned gradient descent (APGD) algorithm, which alternately updates the two factor matrices, eliminating the need for the damping parameter λ and enabling faster convergence with larger step sizes. We theoretically prove that APGD convergences to a near-optimal error at a linear rate. We further show that APGD can be extended to deal with other low-rank matrix estimation tasks, also with a theoretical guarantee of linear convergence. To validate the effectiveness and scalability of the proposed APGD, we conduct simulated and real-world experiments on a wide range of low-rank estimation problems, including noisy matrix sensing, weighted PCA, 1-bit matrix completion, and matrix completion. The extensive results demonstrate that APGD consistently achieves the fastest convergence and the lowest computation time compared to the existing alternatives.
Large volumes of online impressions are sold daily via real-time auctions to deliver targeted advertisements to consumers. Advertisers use data to learn about user preferences and select the most appropriate ad for each user, which also helps them optimize their bids in an ad auction. Although ad exchanges may provide some user data to advertisers, they are usually limited, and advertisers often acquire data from various sources to improve targeting performance. The acquisition of such data can significantly influence the revenue of the ad exchange, which motivates ad exchanges to take actions that reduce advertisers' data acquisition costs and encourage them to buy data. Previous studies have examined the impact of ad exchanges revealing their data to advertisers, but little attention has been paid to the impact of ad exchanges subsidizing advertisers to acquire data from third parties. To address this gap, we propose three subsidy frameworks to increase ad exchange revenue by inducing more advertisers to acquire data: all subsidized (AS), winner subsidized (WS), and loser subsidized. Using a stylized model, we analyze the impact of subsidy provisions on the platform's net revenue. Our results show that WS can be better or worse than AS depending on the cost of data acquisition, its beneficial impact on ad selection, and the distribution of impression values.
Although robust tensor completion has been extensively studied, the effect of incorporating side information has not been explored. In this article, we fill this gap by developing a novel high-order robust tensor completion model that incorporates both latent and explicit side information. We base our model on the transformed t-product because the corresponding tensor tubal rank can characterize the inherent low-rank structure of a tensor. We study the effect of side information on sample complexity and prove that our model needs fewer observations than other tensor recovery methods when side information is perfect. This theoretically shows that informative side information is beneficial for learning. Extensive experimental results on synthetic and real data further demonstrate the superiority of the proposed method over several popular alternatives. In particular, we evaluate the performance of our solution based on two important applications, namely, link prediction in signed networks and rating prediction in recommender systems. We show that the proposed model, which manages to exploit side information in learning, outperforms other methods in the learning of such low-rank tensor data. Furthermore, when dealing with varying dimensions, we also design an online robust tensor completion with side information algorithm and validate its effectiveness using a real-world traffic dataset in the supplementary material.
Federated learning (FL), a distributed learning paradigm focused on preserving data privacy, faces challenges due to varying data distributions among clients, impacting global model performance. To mitigate data heterogeneity, we propose FedUB—a personalized FL framework leveraging uniform feature representation and balancing personalization and collaboration in the classifier. Specifically, the uniform representation (UR) in FedUB provides all clients with a shared feature extractor and a common representation centroid (RC). Achieving this uniformity involves incorporating a regularization term to reduce the gap between global and local RCs. Additionally, an importance estimation of the parameters in the classifier is provided to partition the parameters into two parts: the personalized component and the collaborated component. Specifically, the personalized component adapts to local data, while the collaborated component prevents the classifier from overfitting local data. Theoretically, we establish the existence of the UR, demonstrating its effectiveness in reducing the average generalization bound. Experiments on benchmark datasets consistently demonstrate the performance gains and improved generalization behavior of FedUB.
In this paper, we consider the following dynamic pricing problem. Suppose the market price vt of an item arriving at time t is determined by vt=θTxt, where xt is the feature vector of that item and θ is an unknown vector parameter. The seller has to post prices without knowing θ such that the total regret in time span T is minimized. Considering real-world scenarios in which people may negotiate prices, we propose a model called Second Chance Pricing, in which a seller has a second opportunity to post a price after the first offer is declined. Theoretical analysis shows that a second chance of pricing results in a total regret between O(lnTnlnn+1n) and O(n2lnT), where n is the dimension of the feature space. Experiments on both synthetic data and real data demonstrate significant benefits brought about by the second chance where the regret is only 13% of that of one chance.
In federated learning (FL), model aggregation is a critical step by which multiple clients share their knowledge with one another. However, it is also widely recognized that the aggregated model, when sent back to each client, performs poorly on local data until after several rounds of local training. This temporary performance drop can potentially slow down the convergence of the FL model. Most research in FL regards this performance drop as an inherent cost of knowledge sharing among clients and does not give it special attention. While some studies directly focus on designing techniques to alleviate the issue, an in-depth investigation of the reasons behind this performance drop has yet to be conducted.To address this gap, we conduct a layer-peeled analysis of model aggregation across various datasets and model architectures. Our findings reveal that the performance drop can be attributed to two major consequences of the aggregation process: (1) it disrupts feature variability suppression in deep neural networks (DNNs), and (2) it weakens the coupling between features and subsequent parameters.Based on these findings, we propose several simple yet effective strategies to mitigate the negative impacts of model aggregation while still enjoying the benefit it brings. To the best of our knowledge, our work is the first to conduct a layer-peeled analysis of model aggregation, potentially paving the way for the development of more effective FL algorithms.