A deep hashing method using a time sequence speech visualization feature was proposed to improve speech retrieval accuracy, the efficiency of deep learning, and the noise robustness of speech deep hashing. The speech data is transformed into spectrogram time sequences. Two deep learning models, the 3D CNN (Convolutional Neural Networks)-BiLSTM (Bidirectional Long Short-Term Memory) deep model, and the bidirectional ConvLSTM deep model, are constructed to learn from the time sequence speech visualization feature to generate deep hashing for speech full-text retrieval. Experimental results demonstrate that the proposed models only require fewer iterations to achieve satisfied training accuracy. Moreover, the time sequence speech visualization feature can effectively represent speech content to achieve high retrieval accuracy and recall. Additionally, the proposed deep hashing exhibits high robustness to noise. The proposed method has higher accuracy and is robust to speaker identity compared to existing methods.