WASP PhD students and postdocs from the cluster Sequential decision making and reinforcement learning visited Mila – Quebec AI Institute to exchange ideas and take part in a conference. The trip turned out to be a success when it offered new research perspectives and opportunities to build international connections. Mila is one of WASP’s partner universities, and the visit contributed to strengthening the long-term collaboration between the two research communities.
The visit was organized in August 2026 to strengthen connections between researchers in the WASP cluster and the reinforcement learning community in Montreal. The program combined research presentations, meetings with faculty and PhD students, informal scientific discussions, and participation in Reinforcement Learning Conference 2026.
“The study trip provided valuable opportunities for scientific exchange and networking. Discussions with researchers in Montreal gave the participants feedback and new perspectives on their research, while presentations from the host groups offered insight into current directions in reinforcement learning and machine learning,” says Stefan Stojanovic, the main organizer and PhD student at KTH Royal Institute of Technology.
Research exchange at Mila
During the visit, the group consisting of 11 WASP PhD students and postdocs, met with researchers and PhD students working on reinforcement learning, machine learning, robotics, and related areas. The first part of the program included a meeting with Professor Alex Hernandez-Garcia, who presented his research on machine learning for scientific discovery and introduced GFlowNets, a method with similarities to reinforcement learning that is designed to sample diverse outcomes in proportion to their reward.
The group also met PhD students supervised by Professor Glen Berseth, who co-directs the Robotics and Embodied AI Lab at Mila, and students from Professor Pierre-Luc Bacon’s research group. Their presentations covered topics such as world models, scaling robotic pretraining, adaptive policy priors, and concentration of cumulative rewards in MDPs.
The breadth of topics aligned well with the WASP cluster, which brings together researchers working across several areas of reinforcement learning and sequential decision making.
On Friday, the group joined a larger reinforcement learning meeting organized by Mila. Six WASP PhD students presented their research and received feedback from researchers at Mila and other visiting researchers. The meeting also featured presentations from researchers visiting from institutions including ETH, DeepMind, and the Max Planck Institute.
Perspectives from Reinforcement Learning Conference 2026
The group also participated in Reinforcement Learning Conference 2026, where they attended talks and workshops, presented their work, and connected with the broader international reinforcement learning community.
For participant Jenni Reuben, Industrial Postdoc at KTH Royal Institute of Technology and Research Scientist at Saab Aeronautics, the conference and study trip highlighted several important questions for safe and trustworthy reinforcement learning.
“One key takeaway was that an agent may perform well under normal conditions but still fail when the environment changes. Many discussions therefore focused on robustness, distribution shifts, and how to identify when an agent moves beyond its area of competence,” she says.
She also noted that uncertainty detection becomes valuable only when it leads to an action, such as slowing down, abstaining, transferring control, or activating a safety filter.
“This connection between detecting uncertainty and deciding how to intervene was especially relevant to my own research,” says Reuben.
Reuben also appreciated the format of the conference, where accepted papers were first presented in short oral sessions and then discussed in poster sessions.
“The format made it easier to identify the papers most relevant to me and then follow up with one-to-one discussions with the authors,” she says.
The study trip offered participants new research perspectives, feedback on their own work, and opportunities to build international connections in reinforcement learning, safe autonomous systems, and related areas.

Interested in organizing your own study trip?
Several options are available for PhD students interested in organizing a study trip. Trips can be arranged through a cluster or organized independently as a self-arranged study trip.








Published: September 10th, 2026
[addtoany]