
We are pleased to announce the release of 11 new PLLuM models. Their primary advantage lies in their exceptional proficiency in the Polish language – including official/administrative styles – as well as a deep understanding of native cultural, historical, and legal contexts. These models are designed to support public administration, businesses, and individual users. Crucially, they have been released under open licenses that are fully compliant with the requirements of the EU AI Act.
The Specificity of the New PLLuM Models
The new PLLuM variants can significantly enhance the efficiency of public administration. They are capable of generating texts in over 20 types of official documents, supporting office and operational tasks, interpreting the context of administrative procedures, simplifying complex legal language, and working with standardized legal document templates.
Based on an analysis of real user interactions with PLLuM Chat, we have also developed mechanisms that enable the generation of safer and more precise responses.
Four Model Sizes
The new model family includes four sizes: refreshed versions of 8B, 12B, and 70B, along with a brand-new 4B category:
- 4B – The smallest and fastest models with low computational requirements, ideal for task-specific fine-tuning.
- 8B and 12B – Providing an excellent balance between speed and quality; recommended for production deployments, such as serving as the engine for RAG (Retrieval-Augmented Generation) systems.
- 70B – The largest and most advanced model, designed to handle complex tasks effectively without the need for additional fine-tuning.
All versions are available under open licenses with full documentation compliant with the AI Act, including detailed descriptions of the models, data sources, training methods, and quality evaluation metrics.
Model Training
The models were developed in 2025 commissioned by the Ministry of Digital Affairs as part of the HIVE AI project. The project was implemented by a consortium consisting of: NASK PBI (leader), ACK Cyfronet AGH, Centre for Information Technology (COI), Institute of Computer Science PAS, Institute of Slavic Studies PAS, OPI PIB, Wrocław University of Science and Technology, and the University of Łódź.
The training process was based on a new, rich, and diverse corpus of text materials. The data was collected legally through licensing agreements, public domain sources, and Creative Commons resources.
Representing the Institute of Slavic Studies PAS, the project was coordinated by Dr. hab. Roman Roszko, Prof. IS PAS, with a team including Mgr. Tomasz Bernaś and Mgr. Valéry Trân Thiên, bridging the fields of computer science and linguistics.
Information about the premiere of new models is also available on the website of the Ministry of Digital Affairs.


