
The European Commission has launched a new multilingual AI model, the EU-Institutional-LLM, designed to support all 24 official EU languages effectively.
This open large language model, developed by the Directorate-General for Translation (DGT), aims to provide balanced performance across the diverse linguistic landscape of the European Union. Unlike many existing models that excel in only a few languages, the EU-Institutional-LLM is tailored to understand the specific legislative and policy contexts of the EU, utilizing high-quality multilingual datasets from the DGT.
The model features a base version with approximately 47 billion parameters and employs a Mixture-of-Experts architecture to enhance multilingual capabilities while minimizing issues like catastrophic forgetting. An instruct version further refines the model through supervised fine-tuning and preference alignment, ensuring it meets user needs effectively. The training was conducted using advanced European supercomputing resources, including systems in Luxembourg, Italy, and Spain.
This initiative is significant as it addresses the challenges of multilingual AI, which often favors a limited number of languages. By focusing on all official EU languages and embedding relevant EU knowledge, the Commission aims to enhance AI accessibility in public sectors. Furthermore, this effort reflects a broader European strategy to foster digital sovereignty, emphasizing the importance of developing open AI infrastructures that align with European values and regulatory standards.