We develop a semi-amortized learning framework for downlink beamforming in large-scale sparse multiple-input multiple-output (MIMO) channels. The core of the proposed approach is a deep semi-amortized encoder-decoder network (SA-EDN) architecture composed of three modules: 1) An encoder neural network (NN), deployed at each user, compresses the estimated downlink channel into a low-dimensional latent representation, which is then fed back to the base station (BS); 2) A beamformer decoder NN at the BS, first maps the recovered latent vectors to transmit beamformers and then performs a few steps of gradient ascent to refine the beamformers; and 3) A channel decoder NN, also located at the BS, reconstructs the downlink channels from the recovered latent vectors. The training of SA-EDN leverages two key strategies: 1) A two-phase training scheme, in which the encoder NN and beamformer decoder NN are alternately trained in the first phase, followed by supervised training of the channel decoder NN in the second phase; and 2) Knowledge distillation, where the first training phase starts from supervised training with linear minimum mean-square error (LMMSE) beamformers as labels, and gradually shifts toward unsupervised training using the sum-rate objective. The proposed SA-EDN beamforming framework is extended to both far-field and near-field hybrid beamforming scenarios. Extensive simulation results demonstrate its effectiveness across diverse network and channel conditions, as well as its superiority over several baseline methods.