Skip to content

WASP cluster members in front of the Mila building in Montreal.

WASP PhD students and postdocs from the cluster Sequential decision making and reinforcement learning visited Mila – Quebec AI Institute to exchange ideas and take part in a conference. The trip turned out to be a success when it offered new research perspectives and opportunities to build international connections. Mila is one of WASP’s partner universities, and the visit contributed to strengthening the long-term collaboration between the two research communities.

The visit was organized in August 2026 to strengthen connections between researchers in the WASP cluster and the reinforcement learning community in Montreal. The program combined research presentations, meetings with faculty and PhD students, informal scientific discussions, and participation in Reinforcement Learning Conference 2026.

“The study trip provided valuable opportunities for scientific exchange and networking. Discussions with researchers in Montreal gave the participants feedback and new perspectives on their research, while presentations from the host groups offered insight into current directions in reinforcement learning and machine learning,” says Stefan Stojanovic, the main organizer and PhD student at KTH Royal Institute of Technology.

Research exchange at Mila

During the visit, the group consisting of 11 WASP PhD students and postdocs, met with researchers and PhD students working on reinforcement learning, machine learning, robotics, and related areas. The first part of the program included a meeting with Professor Alex Hernandez-Garcia, who presented his research on machine learning for scientific discovery and introduced GFlowNets, a method with similarities to reinforcement learning that is designed to sample diverse outcomes in proportion to their reward.

The group also met PhD students supervised by Professor Glen Berseth, who co-directs the Robotics and Embodied AI Lab at Mila, and students from Professor Pierre-Luc Bacon’s research group. Their presentations covered topics such as world models, scaling robotic pretraining, adaptive policy priors, and concentration of cumulative rewards in MDPs.

The breadth of topics aligned well with the WASP cluster, which brings together researchers working across several areas of reinforcement learning and sequential decision making.

On Friday, the group joined a larger reinforcement learning meeting organized by Mila. Six WASP PhD students presented their research and received feedback from researchers at Mila and other visiting researchers. The meeting also featured presentations from researchers visiting from institutions including ETH, DeepMind, and the Max Planck Institute.

Perspectives from Reinforcement Learning Conference 2026

The group also participated in Reinforcement Learning Conference 2026, where they attended talks and workshops, presented their work, and connected with the broader international reinforcement learning community.

For participant Jenni Reuben, Industrial Postdoc at KTH Royal Institute of Technology and Research Scientist at Saab Aeronautics, the conference and study trip highlighted several important questions for safe and trustworthy reinforcement learning.

“One key takeaway was that an agent may perform well under normal conditions but still fail when the environment changes. Many discussions therefore focused on robustness, distribution shifts, and how to identify when an agent moves beyond its area of competence,” she says.

She also noted that uncertainty detection becomes valuable only when it leads to an action, such as slowing down, abstaining, transferring control, or activating a safety filter.

“This connection between detecting uncertainty and deciding how to intervene was especially relevant to my own research,” says Reuben.

Reuben also appreciated the format of the conference, where accepted papers were first presented in short oral sessions and then discussed in poster sessions.

“The format made it easier to identify the papers most relevant to me and then follow up with one-to-one discussions with the authors,” she says.

The study trip offered participants new research perspectives, feedback on their own work, and opportunities to build international connections in reinforcement learning, safe autonomous systems, and related areas.

Jenni Reuben
Jenni Reuben, Industrial Postdoc at KTH Royal Institute of Technology and Research Scientist at Saab Aeronautics.

Interested in organizing your own study trip?

Several options are available for PhD students interested in organizing a study trip. Trips can be arranged through a cluster or organized independently as a self-arranged study trip.

Daniel Lawson, PhD from the REAL group at Mila, presenting his work during meeting with WASP visitors.
Daniel Lawson, PhD from the REAL group at Mila, presenting his work during meeting with WASP visitors.
Raghav Bongole, KTH, presenting work at RL group meeting.
Raghav Bongole, KTH, presenting work at RL group meeting.
Jack Sandberg, Mila cluster trip
Jack Sandberg, Chalmers, presenting work at RL group meeting.
Ahmet Balcioglu, Chalmers
Ahmet Balciouglu, Chalmers, presenting work at RL group meeting.
Gabriele Calzolari, Luleå University
Gabriele Calzolari, Luleå University of Technology, presenting work at RL group meeting.
Mika Persson, Chalmers, presenting work at RL group meeting.
Mika Persson, Chalmers, presenting work at RL group meeting.
David Abel, Deepmind, RLC
David Abel, Deepmind, gave a talk titled “Where is learning?” during the Continual RL workshop at RLC 2026.
All attendees of RLC 2026 had the opportunity to enjoy Echo, a spectacular show by Montreal’s Cirque du Soleil.
All attendees of RLC 2026 had the opportunity to enjoy Echo, a spectacular show by Montreal’s Cirque du Soleil.


Published: September 10th, 2026

[addtoany]
We use cookies to personalise content and ads, to provide social media features and to analyse our traffic. We also share information about your use of our site with our social media, advertising and analytics partners. View more
Cookies settings
Accept
Privacy & Cookie policy
Privacy & Cookies policy
Cookie name Active
The WASP website wasp-sweden.org uses cookies. Cookies are small text files that are stored on a visitor’s computer and can be used to follow the visitor’s actions on the website. There are two types of cookie:
  • permanent cookies, which remain on a visitor’s computer for a certain, pre-determined duration,
  • session cookies, which are stored temporarily in the computer memory during the period under which a visitor views the website. Session cookies disappear when the visitor closes the web browser.
Permanent cookies are used to store any personal settings that are used. If you do not want cookies to be used, you can switch them off in the security settings of the web browser. It is also possible to set the security of the web browser such that the computer asks you each time a website wants to store a cookie on your computer. The web browser can also delete previously stored cookies: the help function for the web browser contains more information about this. The Swedish Post and Telecom Authority is the supervisory authority in this field. It provides further information about cookies on its website, www.pts.se.
Save settings
Cookies settings