Sleep is essential for health, yet polysomnography (PSG), the clinical gold standard for comprehensive sleep assessment, is intrusive and impractical for routine use. Sleep sounds offer a non-intrusive alternative, but acoustic signals alone may provide insufficient information for reliable assessment. To address this limitation, we propose a multimodal knowledge distillation (MKD) framework for estimating subjective sleep quality, defined in this study as participants' self-reported satisfaction with each night's sleep. The framework uses richer multimodal information only during training and deploys a sound-based student model at inference time. The teacher model combines sleep-stage features, PSG features, subjective factors, sound features, and demographic factors using hierarchical gated variable selection networks (GVSNs). The student model is restricted to sound features and demographic factors, and response-based distillation transfers the teacher's multimodal decision behavior to this practical input setting. Using 198 nights from 101 adults, we evaluate the framework under both night-wise and subject-wise protocols. Distillation improves the student over the non-distilled model in the repeated splits, with the larger observed gain occurring when test participants are unseen during training. Further analyses show that the preferred distillation weight, the teacher's learned modality weights, and the effect of removing each teacher modality differ between the evaluation protocols. In particular, the ablation results show that a modality's effect on standalone teacher performance does not necessarily match its effect on the distilled student's performance. These findings provide preliminary, dataset-specific evidence of multimodal distillation but do not establish clinical utility or the transfer of any specific modality.