Traditional energy forecasting solutions rely on task-specific supervision and energy asset representations, limiting transferability and the ability to capture general temporal dynamics across heterogeneous assets. We address this by proposing a distributed Joint Embedding Predictive Architecture (JEPA) for self-supervised learning from heterogeneous energy time-series. The framework predicts latent representations of masked temporal segments while integrating temporal observations and contextual information within a shared embedding space. To prevent representation collapse, training combines a latent-space predictive objective with covariance and temporal variance regularization. The evaluation was conducted on energy consumption and generation datasets under data-degradation scenarios and compared with a Transformer forecasting baseline. The learned representations remained stable (cosine similarity ≈0.98; effective rank 185–235). JEPA achieved performance comparable to a Transformer on building energy data, higher R² in 3/5 consumer clusters, and outperformed the baseline on 9/10 unseen PVs (R²=0.73–0.88 vs. <0.45), while showing greater robustness to missing data.