Third-party large language models often lack the domain expertise, language coverage, and regulatory compliance that EU institutions require. We performed continual pretraining of Mixtral Mixture-of-Experts 8 × 7B model on over 100 billion tokens across all 24 official EU languages from EURAMIS, the European Commission’s translation support system. To mitigate catastrophic forgetting, we applied hierarchical data packaging with token density normalization, integrated replay data using language-specific ratios, and employed warm-up cycles with adaptive learning rate scheduling. We evaluated on standard benchmarks, a custom EU Formal Language sentence completion task, and toxicity probes. The results revealed trade-offs between multilingual adaptation and high-resource language preservation. Replay mechanisms partially mitigate degradation. This paper provides empirically grounded guidance for practitioners developing multilingual LLMs for public sector applications using European HPC infrastructure.
更多
查看译文
关键词
Large Language Models,Continual Pretraining,Catastrophic Forgetting,Low-resource Languages