Publication:

Robust Offline Reinforcement Learning for Autonomous Vessel Navigation

 
cris.virtual.department#PLACEHOLDER_PARENT_METADATA_VALUE#
cris.virtual.department#PLACEHOLDER_PARENT_METADATA_VALUE#
cris.virtual.department#PLACEHOLDER_PARENT_METADATA_VALUE#
cris.virtual.orcid0000-0001-9355-6566
cris.virtual.orcid0000-0002-5523-0634
cris.virtual.orcid0009-0000-9395-9437
cris.virtualsource.department1cf77b59-f7f6-4d1d-af45-e08f88df7d20
cris.virtualsource.departmentf790b071-ce23-4ac6-8ece-af46054a6e2c
cris.virtualsource.departmentb628eb4d-93a6-4d76-8c2b-f26c62ac42ce
cris.virtualsource.orcid1cf77b59-f7f6-4d1d-af45-e08f88df7d20
cris.virtualsource.orcidf790b071-ce23-4ac6-8ece-af46054a6e2c
cris.virtualsource.orcidb628eb4d-93a6-4d76-8c2b-f26c62ac42ce
dc.contributor.authorLesy, Bavo
dc.contributor.authorMercelis, Siegfried
dc.contributor.authorAnwar, Ali
dc.date.accessioned2026-08-24T13:37:21Z
dc.date.available2026-08-24T13:37:21Z
dc.date.createdwos2026
dc.date.issued2025
dc.description.abstractReinforcement Learning (RL) has shown promising advances in autonomous navigation, particularly in learning adaptive control policies through interaction with dynamic environments. However, training RL agents typically require direct interaction with the real environment, which can be dangerous and costly. To mitigate this, offline RL has emerged as a promising alternative, where agents learn from pre-collected datasets. In this work, we study the application of offline RL approaches in inland shipping, comparing the performance and robustness of Behavior Cloning (BC) and Conservative Q-Learning (CQL), two offline approaches, with Soft Actor-Critic (SAC), an online RL algorithm. We design experiments to test the robustness of these algorithms within the domain of unmanned surface vehicles (USVs), using an autonomous shipping simulator based on MOOS-IvP with varied environmental conditions. Our results demonstrate that both offline approaches can achieve performance comparable to the expert online policy in nominal conditions while maintaining similar robustness characteristics in out-of-distribution scenarios. Furthermore, we investigate the impact of dataset quality on offline RL performance, showing that with sufficient data, both BC and CQL can learn policies that match the dataset quality and achieve robustness similar to online RL algorithms.
dc.description.wosFundingTextThis work is conducted within the DEFRA AHOI project, funded by the Belgian Royal Higher Institute for Defence, under contract number 23DEFRA002.
dc.identifier.doi10.1088/1742-6596/3123/1/012006
dc.identifier.issn1742-6588
dc.identifier.urihttps://imec-publications.be/handle/20.500.12860/60096
dc.language.isoeng
dc.provenance.editstepusergreet.vanhoof@imec.be
dc.publisherIOP PUBLISHING LTD
dc.source.beginpage012006
dc.source.conference8 th International Conference on Maritime Autonomous Surface Ships (ICMASS) & Intelligent and Smart Shipping Symposium (ISSS)
dc.source.conferencedate2025-10-08
dc.source.conferencelocationHamburg
dc.source.journal8TH INTERNATIONAL CONFERENCE ON MARITIME AUTONOMOUS SURFACE SHIPS, ICMASS 2025 & INTELLIGENT AND SMART SHIPPING SYMPOSIUM, ISSS
dc.source.numberofpages12
dc.title

Robust Offline Reinforcement Learning for Autonomous Vessel Navigation

dc.typeProceedings paper
dspace.entity.typePublication
imec.internal.crawledAt2026-07-14
imec.internal.sourcecrawler
imec.internal.wosCreatedAt2026-07-14
Files

Original bundle

Name:
Lesy_2025_J._Phys.__Conf._Ser._3123_012006 (1).pdf
Size:
726.42 KB
Format:
Adobe Portable Document Format
Description:
Published
Name:
Lesy_2025_J._Phys.__Conf._Ser._3123_012006 (1).pdf
Size:
726.42 KB
Format:
Adobe Portable Document Format
Publication available in collections: