Multimodal sentiment analysis is vulnerable to unstable local cues and temporal asynchrony among textual, visual, and acoustic signals. Existing fusion methods extensively model cross-modal interaction, but lightweight mechanisms that regularize intermediate representations before temporal alignment remain limited. This letter proposes TV-CMF, a time-varying mediation inspired fusion framework for perturbation-stable multimodal sentiment analysis. TV-CMF establishes a constrained routing pathway in which modality-specific temporal states are first reconstructed within a prototype regularized space and subsequently aligned through learnable continuous offsets. A temporal predictability regularizer further constrains the evolution of the routed representations. Here, “mediation-inspired” denotes an enforced representation-routing bottleneck rather than identifiable causal mediation. Experiments on CMU-MOSI, CMU-MOSEI, and perturbations controlled demonstrate visual/acoustic competitive clean-set performance, higher absolute perturbed-set accuracy, and consistent component-wise gains.