Abstract In urban canyons, multipath and non-line-of-sight reception introduce substantial pseudorange biases that degrade smartphone Global Navigation Satellite System (GNSS) positioning. To improve practical training efficiency beyond the multilayer perceptron (MLP)-based PrNet, we develop a hybrid model combining a convolutional neural network (CNN) and a long short-term memory (LSTM) network. The CNN extracts local features from GNSS observations, while the LSTM models temporal dependencies in pseudorange errors. The CNN-LSTM model was evaluated on the Google Smartphone Decimeter Challenge 2021 (GSDC 2021) dataset against weighted least squares (WLS), PrNet, CNN, and LSTM under consistent conditions. Autocorrelation and ablation analyses indicated temporal dependence in pseudorange residuals and complementary contributions from the CNN and LSTM components. Compared with PrNet, CNN-LSTM reduced the median epoch time by 30.1%, increased training throughput by 43.1%, and reduced graphics processing unit (GPU) training-step latency by 33.9%. For the overall trajectory, CNN-LSTM achieved a horizontal positioning root mean square error (RMSE) of 4.638 m, 5.40% lower than PrNet, while reducing the mean absolute error (MAE) and GSDC score by 9.39% and 5.43%, respectively. It also achieved the lowest values for all three metrics in urban interchange, building-obstructed, and open-sky environments, with improvements of 53.13%–73.96% over WLS. These results show that CNN-LSTM provides a better balance between positioning accuracy and practical GPU training efficiency than PrNet under the evaluated conditions.