
Computed tomography (CT) is a significant clinical detection method without invasive procedures. However, high-dose radiation is associated with elevated risk of cancer. Low-dose CT reconstruction techniques have the potential to significantly reduce this risk. These techniques reduce radiation dose by two means: low radiation intensity and incomplete scans. The focus of this review is incomplete-scan CT reconstruction techniques, including sparse-view and limited-angle CT reconstruction. In the context of incomplete scans, CT images obtained by conventional CT reconstruction techniques exhibit noises and streak artifacts. Deep learning techniques are changing this situation. However, there are few reviews that recently focus on the fast developing area. This paper reviews recent research on CT reconstruction of incomplete scans via deep learning and classifies them into two categories including image domain and reconstruction domain processing methods. This study also evaluates the strengths and limitations of these approaches and discusses future prospects, highlighting promising technologies and directions.
Reconfigurable Intelligent Surfaces (RISs) have emerged as a promising technology for enhancing wireless communication by reconfiguring the propagation environment. Most existing studies assume ideal RIS elements with zero electrical resistance, which limits their applicability in realistic settings. This study considers a practical RIS-aided wideband OFDM system with resistive reflecting elements (REs), where a single-antenna source communicates with a single-antenna destination. The goal is to maximize the transmission rate by jointly optimizing the subcarrier power allocation and RE phase shifts. An alternating optimization (AO) framework is proposed, in which each variable is updated iteratively, and the phase shifts are refined via gradient descent (GD). A joint optimization strategy was also considered for comparison. Furthermore, a low-complexity alternative based on solving a channel power maximization (CPM) problem is introduced to obtain suboptimal phase shifts. The numerical results show that the AO method achieves near-optimal performance with a significantly reduced computation time compared to the successive convex approximation (SCA). Although the GD-based CPM method incurs higher complexity than dimension-wise sinusoidal maximization (DSM), it yields superior rate performance and improved robustness against RE resistance variations, making it more suitable for practical RIS deployment.
Non-autoregressive speech translation (NAR-ST) offers low latency through parallel decoding but often falls short of autoregressive () systems in producing fluent and well-structured translations. The authors present UTRo-NAST, a novel NAR-ST framework that decomposes speech translation into three consecutive subtasks: source speech understanding, word-by-word translation and target-side reordering. This divide-and-conquer design enhances interpretability and training stability while preserving fully parallel inference. To further improve performance, the authors introduce a plug-and-play large language models ()-augmented post-correction strategy that refines UTRo-NAST outputs with prompting mechanisms. Experiments on the MuST-C benchmark across eight language pairs show that UTRo-NAST consistently outperforms existing NAR-ST models and delivers translation quality comparable to strong baselines, while maintaining faster decoding. The LLM-augmented variant achieves competitive or superior results compared to recent LLM-integrated systems, offering a practical path to stronger NAR-ST without iterative refinement or costly fine-tuning. Overall, these results demonstrate the effectiveness and scalability of UTRo-NAST for practical speech translation.
Automatic speech recognition (ASR) provides a scalable solution for annotating large speech datasets, yet inherent transcription errors significantly complicate text-to-speech (TTS) training. While modern diffusion and flow-based architectures achieve high-quality generation, their performance under noisy transcription conditions remains underexplored. This study aims to investigate the robustness of flow- (VITS) and diffusion-based (GradTTS, DVT) models trained on simulated noisy transcriptions, using the autoregressive Tacotron 2 as a baseline. The authors train models on data sets with varying noise levels and evaluate the resulting speech quality using both objective intelligibility metrics and subjective naturalness ratings. Experimental results demonstrate that diffusion-based models exhibit superior robustness, maintaining high intelligibility and naturalness even under high-noise conditions. Furthermore, the behavioral analysis reveals a latent domain separation phenomenon: noisy models spontaneously organize text representations based on transcription quality, despite the absence of explicit labels during training. The authors find that this separation correlates with the resulting text features and degraded synthesis performance. To mitigate this degradation, the authors investigate a text prompt strategy that prepends reliably synthesizable text fragments to guide the model toward activating higher-quality representations. This lightweight approach improves synthesis stability without requiring model fine-tuning.
In recent years, few studies have been conducted on the hierarchical assisted diagnosis of cervical lesions and most of them are still insufficient and their clinical significance is extremely limited. Based on this, this paper designs densely contextual transformer networks (DCTNet) to diagnose cervical lesions. DCTNet benefits from its internal DCT module and multi-spectral channel attention (MSCA) module to diagnose cervical lesions efficiently. First, the DCT module effectively combines the ability of the transformer to capture global contextual information with the ability of the convolutional neural networks to capture neighboring local contextual information, which enhances the visual representation of the network. Second, the MSCA module can effectively use the information of more frequency components to enrich the feature representation ability of the network. The results show that under the same research conditions, the diagnostic performance of DCTNet is generally superior to other current diagnostic methods; the area under the curve value is 0.8332. While the clinical diagnostic performance reaches the level of advanced physicians, close to the expert physicians, with excellent auxiliary diagnostic performance. Therefore, DCTNet can effectively alleviate the pressure on physicians in resource-poor areas.
Demosaicing is crucial in digital imaging, reconstructing full-color images from partial red, green and blue sensor data with two-thirds of the pixel information missing. Although deep learning has improved performance, its large models and high computational demands are unsuitable for resource-limited edge devices. This work presents a green U-shaped image demosaicing (GUSID) method, a novel approach grounded in green learning (GL) principles. GUSID addresses the limitations of traditional methods by offering a lightweight, transparent and efficient solution. Unlike neural network-based methods, GUSID completely avoids deep learning. Instead, GUSID uses unsupervised representation learning for effective feature extraction and supervised feature learning to enhance computational efficiency and ensure high-quality performance. GUSID's compact design minimizes computational overhead while maintaining competitive accuracy. Its support for parallelized training also enables fast execution, making it an ideal choice for real-time vision applications on resource-constrained devices.
Estimating individualized treatment effects (ITE) is critical for personalized medicine, yet it remains a challenge due to retrospective observational data, which suffer from selection bias in clinical practice and the complexity of multimodal data used for patient status depiction. In this work, the authors develop an end-to-end deep learning (DL) framework that incorporates multimodal patient data and multiple treatments for accurate ITE inference in a retrospective head and neck cancer (HNC) study. A possible solution is concatenating the factors and adapting adversarial training, which has shown great promise on tabular data, to disentangle patient characteristics from patient status features to mitigate treatment selection bias. However, this approach suffers from instability when applied to complex multimodal patient data and multiple treatment options. For flexible and efficient treatment-conditioned information fusion, they propose a bi-stage adaptive instance normalization (Bi-AdaIN) to inject relevant factors into corresponding layers, an approach that is also robust to missing values. Furthermore, they propose to disentangle status features from the multi-treatment variable using mutual information (MI) regularization, enabling more accurate predictions of patient-specific outcomes for both factual and counterfactual data. The authors evaluated their model on the RADCURE dataset, comprising 3,346 HNC cases with CT scans and multiple clinical variables who received radiotherapy or additional chemotherapy and EGFRI. The Bias-Adjusted Treatment Effect (BATE) is substantially reduced compared to the conventional direct ITE method (which does not consider treatment bias) and to adversarial training, indicating a more robust estimation of causal effects. This work is one of the first DL-based studies to address ITE estimation using multimodal medical imaging, offering a promising approach to counterfactual reasoning in clinical oncology for decision support.
This study investigates perceptual similarity at two levels: music tracks (track-level) and the individual instrumental parts that compose them (part-level). A previous work performed a study on perceptual part-level similarity toward developing a model that estimates part-level similarity. An ABX-style listening test with 632 participants was conducted, which evaluated similarity at both levels from the perspectives of timbre, melody, rhythm and overall. Although a previous work contributed some knowledge from the evaluations, further insights are needed to support the development of future estimation models. Specifically, important questions remain regarding the correspondence between track- and part-level similarity, the generalizability of findings across multiple models, and the validity of the conventional learning method in terms of perceptual similarity. This study revealed the following key findings: (1) the instrumental parts that predominantly affect the track-level similarity differ across music triplets and listeners, with the influence of the differences across music triplets exceeding the differences across listeners, indicating that part-level similarity helps in estimating track-level similarity; (2) when a temporal averaging is applied, the output of the deep learning models shows a closer correspondence with the perceptual evaluation based on timbre than on rhythm, indicating a potential area for improvement in the models; (3) the similarity between temporally distinct segments within the same music track is significantly perceived to be significantly higher than that between segments from different tracks, which supports the assumption of the conventional unsupervised learning method developed for music similarity estimation.
Zero-shot speech enhancement (SE) aims to improve speech quality in unseen acoustic conditions without requiring task-specific fine-tuning. This work proposes expanded noise modeling for scalable and adaptive zero-shot speech enhancement (EN-AZS), a zero-shot SE framework with expanded noise modeling, built upon an optimized Undiff-based architecture. By extending the noise model to cover a wider range of acoustic variability and incorporating mechanisms such as speech quality scoring and coefficient calculation with controlled update strategies, EN-AZS effectively enhances speech clarity while avoiding overfitting to observed mixtures. Extensive experiments on TIMIT-N6/N9/N15, VCTK-DM and MUSAN data sets demonstrate that EN-AZS consistently outperforms both supervised and unsupervised baseline methods, particularly under mismatched noise types, speech characteristics and SNR conditions. Ablation studies further validate the importance of the speech quality guidance and coefficient calculation mechanisms and update strategies. EN-AZS provides a scalable, plug-and-play solution for robust zero-shot SE, offering a promising approach for real-world applications with diverse and unpredictable noise conditions.
This paper proposes a semisupervised audio-text contrastive learning method based on pseudo-text inputs. The proposed method converts unlabeled audio into effective training data without requiring additional annotations. Its key idea is to mix a labeled audio clip with an unlabeled one. Because the unlabeled clip lacks a textual counterpart, the authors generate a pseudo-text input for the unlabeled clip through an audio-to-text mapper (a2t). The authors investigate two mapper variants: a simple multi-layer perceptron (MLP) that outputs a single vector and a Transformer decoder that produces a short sequence of query tokens; both are trained jointly with the encoders. Training is driven by three InfoNCE losses: one on labeled pairs, one on mixed (labeled + unlabeled) pairs and one on unlabeled audio-pseudo-text pairs. Gaussian noise is added to the audio embedding to regularize the pseudo-text mapping. On cross-modal retrieval, our method yields a 6.6% relative improvement in Recall@1 on AudioCaps and a 4.4% gain on Clotho over a fine-tuned CLAP baseline. Without requiring any captions for the additional audio, our method surpasses baselines trained on larger fully-labeled data sets.
Adequate hydration is essential for cardiovascular stability, thermoregulation and cognitive performance, yet current assessment methods are invasive or laboratory-dependent, limiting their use for continuous monitoring. This study proposes a noninvasive approach for classifying drinking behavior using photoplethysmography (PPG)-derived waveform features. PPG signals from 155 participants were collected under standardized conditions and grouped by self-reported fluid intake (2-4, 4-6 and >6 cups). Fiducial-point-based timing intervals, amplitude ratios and morphological descriptors were extracted to quantify cardiovascular changes. To address class imbalance, up-sampling, synthetic minority over-sampling technique (SMOTE), generative adversarial network-based feature synthesis with physiological consistency checks and image-based augmentation were applied. Model performance, evaluated using an 80:20 stratified hold-out repeated across ten runs, achieved high accuracy with up-sampling and image-based augmentation (97% and 91%; AUC approximate to 1.00 and 0.99), while SMOTE performed comparably (AUC 0.95). Kruskal-Wallis analysis confirmed significant hydration-related differences in time-domain (e.g. systolic-to-diastolic ratio) and amplitude-domain indices (e.g. pulse amplitude index), with small-to-moderate effect sizes. These results demonstrate that PPG features can capture hydration-related cardiovascular variability and support behavioral classification. The framework offers a foundation for wearable, real-time hydration monitoring and future integration with objective intake measures and multi-site PPG acquisition.
In recent years, automatic speech recognition (ASR) systems have become essential components of various computer applications, enabling seamless communication between humans and machines. N-best reranking and error correction are widely used post-processing techniques, each contributing to improving ASR output with distinct strengths. Building on these advantages, they propose a comprehensive framework designed to achieve superior performance. The framework begins with a text correction module to rectify and expand the original N-best list. This augmented list is then evaluated and reranked jointly by a text rescoring module and a text-speech matching module. The text rescoring module assesses hypotheses from a textual perspective, while the text-speech matching module evaluates each hypothesis from the alignment between speech and text. The authors explore three distinct strategies for implementing the text-speech matching module. Using the publicly available HypR benchmark, they fairly evaluated the proposed framework against other methods. Experimental results demonstrate the feasibility of the proposed framework and highlight the contributions of this study.
The ability of artificial intelligence (AI) models to generate highly realistic human voices has advanced rapidly. These technologies power accessibility tools, virtual assistants and creative applications, but they also enable harmful uses, including impersonation, fraud and disinformation. Recent incidents of voice cloning scams targeting businesses and political leaders underscore the urgent need for robust safeguards. Unlike image and video deepfakes, the detection of synthetic voices poses unique challenges due to the complexity of phonetics, prosody and auditory perception. This survey offers a comprehensive overview of AI voice generation and detection methods, encompassing both the technical foundations and the latest state-of-the-art advances. This study also identifies key open challenges, benchmark resources and future directions to make this survey useful for future researchers.
This paper introduces the physics-inspired synthesized underwater image data set (PHISWID), a data set tailored for enhancing underwater image processing through physics-inspired image synthesis. For underwater image enhancement, data-driven approaches (e.g. deep neural networks) typically demand extensive data sets, yet acquiring paired clean atmospheric images and degraded underwater images poses significant challenges. Existing data sets have limited contributions to image enhancement due to a lack of physics models, publicity and ground-truth atmospheric images. PHISWID addresses these issues by offering a set of paired atmospheric and underwater images. Specifically, underwater images are synthetically degraded by color degradation, haze and marine snow artifacts from atmospheric RGB-D images. It is enabled based on a physics-based underwater image observation model. Their synthetic approach generates a large quantity of the pairs, enabling effective training of deep neural networks and objective image quality assessment. Through benchmark experiments with some data sets and image enhancement methods, they validate that their data set can improve the image enhancement performance. Their data set, which is publicly available, contributes to the development of underwater image processing.
Irregular multivariate time series with missing values present significant challenges for predictive modeling in domains such as healthcare. While deep learning approaches often focus on temporal interpolation or complex architectures to handle irregularities, we propose a simpler yet effective alternative: extracting time-agnostic summary statistics to eliminate the temporal axis. Our method computes four key features per variable - mean and standard deviation of observed values, as well as the mean and variability of changes between consecutive observations - to create a fixed-dimensional representation. These features are then used with standard classifiers, such as logistic regression and XGBoost. Evaluated on four biomedical datasets (PhysioNet Challenge 2012, 2019, PAMAP2 and MIMIC-III), our approach achieves state-of-the-art performance, surpassing recent transformer and graph-based models by 0.5-1.7% in AUROC/AUPRC and 1.1-1.7% in accuracy/F1-score, while reducing computational complexity. Ablation studies demonstrate that feature extraction - not classifier choice - drives performance gains, and our summary statistics outperform raw/imputed input in most benchmarks. In particular, we identify scenarios where missing patterns themselves encode predictive signals, as in sepsis prediction (PhysioNet, 2019), where missing indicators alone can achieve 94.2% AUROC with XGBoost, only 1.6% lower than using original raw data as input. Our results challenge the necessity of complex temporal modeling when task objectives permit time-agnostic representations, providing an efficient and interpretable solution for irregular time series classification.
High-resolution radiography, a noninvasive medical technology, plays a crucial role in diagnosing and determining treatment plans for tumors and other diseases. Super-resolution images recovered from low-resolution samples would be a more economical direction to explore. This paper proposes the depthwise convolution multiscale network transformer (D-MNet Transformer) for super-resolution tasks in radiological images. First, the D-MNet encoder is specifically designed to adapt to the data distribution of radiological images, excelling at detecting subtle alterations within the local areas of the skeletal structure and fusing information at various scales. First, to better adapt to the distribution of radiological images, the D-MNet encoder excels in detecting subtle alterations within the local areas of the skeletal structure and fusing information at various scales. Then, the transformer-based decoder integrates global contextual information for reconstruction. Furthermore, this paper proposes a task-specific pixel-level masking strategy to enhance model performance when training on small data sets. Experimental results demonstrate that this method achieves state-of-the-art performance in PSNR.
In this paper, the authors propose an interpretable denoising method for graph signals using regularization by denoising (). RED is a technique developed for image restoration that uses an efficient (and sometimes black-box) denoiser in the regularization term of the optimization problem. By using RED, optimization problems can be designed with the explicit use of the denoiser, and the gradient of the regularization term can be easily computed under mild conditions. The authors adapt for denoising of graph signals beyond image processing. The authors show that many graph signal denoisers, including graph neural networks, theoretically or practically satisfy the conditions for . The authors also study the effectiveness of from a graph filter perspective. Furthermore, the authors propose supervised and unsupervised parameter estimation methods based on deep algorithm unrolling (). These methods aim to enhance the algorithm applicability, particularly in the unsupervised setting. Denoising experiments for synthetic and real-world data sets show that the proposed method improves signal denoising accuracy in mean squared error compared to existing graph signal denoising methods.
Nuclei segmentation is the cornerstone task in histology image reading, shedding light on the underlying molecular patterns and leading to disease or cancer diagnosis. Yet, it is a laborious task that requires expertise from trained physicians. The large nuclei variability across different organ tissues and acquisition processes challenges the automation of this task. On the other hand, data annotations are expensive to obtain, and thus, Deep Learning (DL) models are challenged to generalize to unseen organs or different domains. This work proposes Local-to-Global NuSegHop (LG-NuSegHop), a self-supervised pipeline developed on prior knowledge of the problem and molecular biology. There are three distinct modules: (1) a set of local processing operations to generate a pseudolabel, (2) NuSegHop a novel data-driven feature extraction model and (3) a set of global operations to post-process the predictions of NuSegHop. Notably, even though the proposed pipeline uses { no manually annotated training data} or domain adaptation, it maintains a good generalization performance on other datasets. Experiments in three publicly available datasets show that our method outperforms other self-supervised and weakly supervised methods while having a competitive standing among fully supervised methods. Remarkably, every module within LG-NuSegHop is transparent and explainable to physicians.