A brain-computer interface (BCI) that decodes speech directly from neural activity provides a rapid and natural means of communication for individuals with speech impairments or aphasia. Recent advances in deep learning have led to several studies demonstrating promising outcomes using electrocorticography (ECoG) placed on cortical surfaces. In contrast, stereo electroencephalography (SEEG) captures neural signals from multiple brain regions, including the cortex and subcortex. These signals, encompassing rich information from deeper brain structures, have significant potential to enhance the characterization of speech generation processes and improve decoding performance. However, effective SEEG-based decoding schemes remain limited. Existing deep learning methods often struggle with overfitting due to insufficient data. In addition, the reconstructed speech tends to be blurry and detail-deficient, indicating an urgent need for more refined SEEG modeling. To address these issues, a convolutional encoder-decoder with scale-recursive reconstructor (ConvED-SR) is proposed for SEEG speech decoding. ConvED-SR first extracts multiscale speech-related features from SEEG signal using a convolutional encoder-decoder architecture. This creates a compact and efficient latent feature space with reduced parameters, thus mitigating overfitting. Furthermore, these multiscale features, effectively characterizing the intricate relationships between deeper neural signals and speech, are used to generate a refined Mel-spectrogram by a scale-recursive reconstructor. The reconstructor initially models low-frequency information, gradually interacts with high-frequency information, and ultimately refines a coarse Mel-spectrogram into a detailed final one. Finally, a HiFi-GAN vocoder converts the spectrogram into speech. Comprehensive experimental results on the SingleWordProductionDutch-iBIDS dataset demonstrate that ConvED-SR achieves superior performance, providing a promising solution for SEEG-based speech decoding.