Project

Uncertainty-Aware Deep Reinforcement Learning

The objective of this project was to improve the data efficiency and safety of deep reinforcement learning (RL) for deployment on physical hardware by means of uncertainty-aware model-based reinforcement learning (MBRL). Here a dynamics model of the environment is trained from observed interaction data and leveraged to predict outcomes in the real world. We focus on avoiding model exploitation, i.e., avoiding to rely on model predictions in areas with low data coverage where the model is unreliable. By pursuing this line of research, we hope to build algorithms that foster efficient and save physical intelligence tailored to the needs of practitioners.

The work consisted of fundamental research in algorithm design, which required a large number of empirical evaluations, each with substantial computational cost. Access to the granted HPC resources was therefore a prerequisite for obtaining the results reported below.

Project Details

Project term

June 27, 2025–July 26, 2026

Affiliations

RWTH Aachen University

Institute

Institute for Data Science in Mechanical Engineering

Principal Investigator

Prof. Dr. Sebastian Trimpe

Methods

As outlined in the project proposal, and building on our prior work in uncertainty-aware MBRL [1, 2], we pursued three directions: (i) application to physical hardware; (ii) transfer to latent world models with visual input; and (iii) use of the uncertainty-aware world model for safe exploration. For (i), we developed a JAX implementation of [2] that learns control behaviors on hardware within minutes and accounts for the partial observability commonly encountered in physical systems. For (ii), we examined the limitations of epistemic uncertainty quantification in latent world models. For (iii), we developed two safe-exploration methods, drawing on viability theory and model predictive safety filtering, respectively.

[1] Frauenknecht, Eisele, Subhasish, Solowjow, Trimpe. Trust the Model Where It Trusts Itself—Model-Based Actor-Critic with Uncertainty-Aware Rollout Adaption. ICML 2024.

[2] Frauenknecht, Subhasish, Solowjow, Trimpe. On Rollouts in Model-Based Reinforcement Learning. ICLR 2025.

Results

(i) Learning on hardware. The JAX implementation reduces training time from hours, as required by prior implementations, to minutes. We demonstrate efficient learning on the Mini-Wheelbot, a challenging benchmark task for robot control. The approach is competitive with common sim-to-real methods while requiring substantially less manual engineering of simulators. These results are under review at IEEE Robotics and Automation Letters.

(ii) Latent world models. We provide, to our knowledge, the first systematic analysis of epistemic uncertainty estimates in latent world models. The result is surprising: current latent world model architectures do not support fine-grained uncertainty quantification. This work has been accepted at the 2026 Conference on Reinforcement Learning.

(iii) Safe exploration. We developed two complementary approaches to safe exploration based on uncertainty-aware MBRL. UPSi constructs a robust MPC tube from a probabilistic neural network dynamics model and introduces model uncertainty as a nonlinear constraint in a model predictive control problem, enabling safe exploration through a predictive safety filter. Dyna-SAuR learns a separating hyperplane that restricts the admissible actions to viable ones which preserve model certainty. UPSi has been accepted at the 2026 Conference on Reinforcement Learning; Dyna-SAuR is under review for the 2026 Conference on Neural Information Processing Systems.

Discussion

Distributing the compute resources among several closely collaborating researchers, each leading a different subtopic, enabled several contributions, reflected in publications at leading venues in the field. The work supported by this project further received two student awards at RWTH Aachen, and four of the master’s students involved have since begun doctoral research in adjacent areas such as dynamics modeling and safe learning. Further, the presented results where a main part of invited keynote presentations by Prof. Trimpe, for example, at the 2nd German Robotics Conference or the 2025 RITA Conference.

Several open questions remain for future work, including the transfer of MBRL to contact-rich tasks such as manipulation, the further development of MBRL-based viability filters, and the construction of latent world models that permit reliable uncertainty quantification.

Additional Project Information

DFG classification: 407-01 Automation, Control Systems, Robotics, Mechatronics, Cyber Physical Systems
Software: Python3
Cluster: CLAIX

Publications

On Rollouts in Model-Based Reinforcement Learning,
Bernd Frauenknecht, Devdutt Subhasish, Friedrich Solowjow, Sebastian Trimpe,
International Conference on Learning Representations, 2025

Uncertainty-Aware Predictive Safety Filter for Probabilistic Neural Network Dynamics,
Bernd Frauenknecht, Lukas Kesper, Daniel Mayfrank, Henrik Hose, Sebastian Trimpe,
Reinforcement Learning Journal Vol.7, 2026

Biased Dreams: Limitations to Epistemic Uncertainty Quantification in Latent Dynamics Models,
Julia Berger, Bernd Frauenknecht, Sebastian Trimpe, Bastian Leibe,
Reinforcement Learning Journal Vol.7, 2026

Dyna-Style Safety Augmented Reinforcement Learning: Staying Safe in the Face of Uncertainty,
Artur Eisele, Bernd Frauenknecht, Friedrich Solowjow, Sebastian Trimpe,
White Paper, 2026

All Models are Wrong, Knowing Where is Useful: On Model Uncertainty in Reinforcement Learning,
Bernd Frauenknecht, Devdutt Subhasish, Artur Eisele, Friedrich Solowjow, Sebastian Trimpe,
Workshop on Uncertainty in Open World Robotics of the 2026 IEEE International Conference on Robotics & Automation, 2026

Learning to Race in Minutes: Infoprop Dyna on the Mini Wheelbot;
Devdutt Subhasish, Henrik Hose, Sebastian Trimpe,
German Robotics Conference Vol.2, 2026

Thesis:
An Information Theoretic Perspective on Synthetic Rollouts in Model-Based Reinforcement Learning,
Devdutt Subhasish,
Master thesis, September 2024

Distributional Model-Based Actor-Critic for Data Efficient and Uncertainty-Aware Reinforcement Learning,
Jonas Hertrampf,
Master thesis, September 2024

Efficient Latent Imagination via Uncertainty-Aware World Models in Deep Reinforcement Learning,
Julia Berger,
Master thesis, January 2025

Inferring Safety Filters for Deep Reinforcement Learning from Uncertainty- Aware Dynamics Models,
Artur Eisele,
Master thesis, May 2025

Reinforcement Learning with Uncertainty-Aware Model Predictive Safety Certification,
Lukas Kesper,
Master thesis, July 2025

Smoothed Discrete Latent Dynamics Models for Reliable Uncertainty Quantification in Model-Based Reinforcement Learning,
Amine Chalghoum,
Master thesis, 2026

Multi-Skill Reinforcement Learning for the Mini-Wheelbot using Linear Temporal Logic,
Jakob Koempel,
Master Thesis, 2026