Developmental voice change exposes a gap in models of self-voice processing: these models explain predictive control and self-voice recognition but not how vocal signals integrate into self-representations. I propose a hierarchical framework in which recursive interactions across sensorimotor predictions, self-voice representations, and higher-order self-representations support the emergence of the vocal self.