The aim of this paper is to investigate the use of multiplewindow Short-Time Fourier Transform (STFT) representation for single sensor source separation. We propose to iteratively split the observed signal into target sources and residuals. Each target source is modeled as the sum of elementary components with known power spectral densities (PSDs). The approach involves a non negative decomposition of the spectra of the observed mixure in a given frame into a dictionary of PSDs. The resolution of the PSDs varies at each iteration of the algorithm. The decomposition into source signals and residual signals employs a confidence measure, which is based on the Fisher information matrix of the expansion coefficients. We demonstrate the improved performance of the proposed method on mixtures of voice and music signals.