Recent single-image super-resolution methods leverage the generative priors of diffusion models to achieve superior perceptual quality in the reconstructed images. However, these perceptual-driven methods suffer from severe fidelity distortion in stereo super-resolution tasks due to the stereo inconsistency introduced by the stochastic nature of the diffusion generation process. We address this challenge by leveraging stereo constraints to guide diffusion at both noise sampling and diffusion feature fusion, allowing the model to handle complex low-resolution degradation while maintaining cross-view consistency, achieving a favorable perceptual-fidelity balance. To be specific, we address the fundamental problem that the inherent stochasticity of diffusion disrupts stereo consistency by proposing a stereo-consistent noise sampling optimization strategy, where we obtain the leftview noise by sampling an approximate optimal initial noise and optimizing the variance noise via a perceptual-consistency loss, while the right-view noise is obtained by disparity-aware warping. This approach drives the diffusion process toward outputs that satisfy stereo high-frequency detail consistency while delivering improved perceptual quality. To enhance stereo consistency at the structural and semantic levels, we design a Disparity-Aware State-Space Module, which performs stereo feature fusion via a sequentially bi-directional cross-scanning approach that is more suitable for the diffusion feature space than traditional attention operation. Together, we build a diffusion-based stereo superresolution framework, DiffSSR+, which progressively integrates stereo guidance into the diffusion process to achieve a better fidelity-perception trade-off, and is plug-and-play to a broad range of diffusion SR architectures.