SE.LLMA (Swedish Large LLM Arena) is a national initiative with the goal of developing, training and making available high-quality open foundation models with hundreds of billions of parameters, adapted to Swedish linguistic characteristics, cultural context and regulatory requirements. The project is part of the Wallenberg AI, Autonomous Systems and Software Program (WASP) and is carried out within WARA AI-TRICS (WASP Research Arena for AI-driven Technologies for Resilient Infrastructure and Critical Systems).
The position is based at NAISS, Sweden’s national infrastructure for high-performance computing, AI and data, providing world-class compute and storage resources. SE.LLMA is conducted in close collaboration between universities, government agencies and industrial partners, and is built on open research, open models and open-source software. This is a unique opportunity to help develop some of Europe’s most advanced open language models from the ground up.
NAISS, the National Academic Infrastructure for Supercomputing in Sweden, provides academic users with high-performance computing resources, storage capacity, and data services. NAISS is hosted by Linköping University and has 12 partner universities across Sweden.
The position
As an Application Expert for AI Infrastructure, you will be part of the team developing the technical software environment for training, evaluating and operating the next generation of open language models. You will be responsible for the software infrastructure required to develop, train, test and deploy models with billions of parameters on some of Europe’s most advanced supercomputers. The duties include distributed AI frameworks and container platforms to communication libraries, storage systems and training pipelines. A central part of the role is developing, configuring and optimizing software stacks for distributed training, inference and evaluation of large language models. You will work with modern open AI frameworks and develop reproducible and portable container-based environments that enable robust and efficient training and evaluation across thousands of accelerators.
Your work will also include identifying and eliminating bottlenecks in computation, communication and storage, optimizing GPU utilization, parallelization, inter-node communication and I/O performance on parallel file systems. You will develop solutions for monitoring, logging, debugging and recovering long-running training jobs, build automated training and evaluation pipelines, and serve as the technical link between AI engineers, researchers, software developers and the NAISS operations team. Together you will build the national AI infrastructure that enables the training and evaluation of some of Europe’s most advanced open language models.
The role involves close collaboration with international research projects, hardware vendors and developers of leading AI frameworks, providing excellent opportunities to influence the future technical platforms for large-scale artificial intelligence.
You will work at the forefront of technology with access to the latest generation of GPU accelerators, high-speed interconnects, parallel file systems and distributed AI frameworks. This position is ideal for someone who wants to combine deep systems expertise with large-scale AI and enjoys solving technical challenges where performance, scalability, robustness and reproducibility are critical.