Single-cell RNA sequencing (scRNA-seq) provides a novel perspective to explore cellular biology at the single-cell resolution. Single-cell clustering is a crucial step to reveal cell types and the corresponding biological functions. However, when dealing with the high dimensionality and complexity of scRNA-seq data, existing deep models fail to comprehensively capture the intrinsic attribute information and structural relationships within the data. In this study, we propose a novel single-cell deep clustering model named scDFVA. The proposed scDFVA consists of a variational graph attention autoencoder (AE), a zero-inflated negative binomial (ZINB) based AE, and a self-supervised clustering. To better simulate sparse and zero-inflated scRNA-seq data, we incorporate the ZINB model into the AE. The variational graph attention AE is introduced to learn the cell structure information. scDFVA achieves representation learning within a joint framework comprising a ZINB-based AE and a variational graph attention AE, effectively fusing gene expression and cell structure information. Furthermore, scDFVA performs self-supervised clustering training on the latent fusion representations of cells to achieve mutual supervision between representation learning and clustering. Experiments indicated that scDFVA outperformed several other competing methods, demonstrating that our method is beneficial in single-cell clustering.