Diffusion Transformers (DiTs) achieve remarkable performance within image generation. Conventionally, DiTs are constructed by stacking serial isotropic transformers, which face significant quadratic computational cost. However, through empirical analysis, we find that DiTs do not rely as heavily on long-distance information as previously believed. In fact, most layers exhibit significant redundancy in long-distance computation. Additionally, conventional attention mechanisms suffer from low-frequency inertia, limiting their efficiency. To address these issues, we propose Pseudo Shifted Window Attention (PSWA), which fundamentally mitigates long-distance attention redundancy. PSWA achieves moderate global-local information through Static Window Attention. It further utilizes a high-frequency bridging branch to enrich the high-frequency information and strengthen inter-window connections. Furthermore, we pioneer the concept of Kth-order Attention and propose the Progressive Coverage Channel Allocation (PCCA) strategy that captures high-order attention by reallocating the existing channel budget. Based on these innovations, we propose a series of Pseudo Progressive Diffusion Transformer (PiT). Extensive experiments show superior performance of PiT; for example, PiT-L achieves 54% FID improvement over DiT-XL/2 with less computation.