Publication:

LiDAR-BIND-T: Temporally Consistent Sensor Modality Translation and Fusion for Robotic Applications

 
cris.virtual.department#PLACEHOLDER_PARENT_METADATA_VALUE#
cris.virtual.department#PLACEHOLDER_PARENT_METADATA_VALUE#
cris.virtual.department#PLACEHOLDER_PARENT_METADATA_VALUE#
cris.virtual.orcid0000-0001-9355-6566
cris.virtual.orcid0000-0002-5523-0634
cris.virtual.orcid0000-0002-4340-9776
cris.virtualsource.department1cf77b59-f7f6-4d1d-af45-e08f88df7d20
cris.virtualsource.departmentf790b071-ce23-4ac6-8ece-af46054a6e2c
cris.virtualsource.department468fc189-f7e0-44c3-a579-522fc479a0d2
cris.virtualsource.orcid1cf77b59-f7f6-4d1d-af45-e08f88df7d20
cris.virtualsource.orcidf790b071-ce23-4ac6-8ece-af46054a6e2c
cris.virtualsource.orcid468fc189-f7e0-44c3-a579-522fc479a0d2
dc.contributor.authorBalemans, Niels
dc.contributor.authorAnwar, Ali
dc.contributor.authorSteckel, Jan
dc.contributor.authorMercelis, Siegfried
dc.date.accessioned2026-09-22T09:49:51Z
dc.date.available2026-09-22T09:49:51Z
dc.date.createdwos2026
dc.date.issued2026
dc.description.abstractRobust autonomous navigation requires reliable perception when optical sensors fail under adverse conditions, such as fog, rain, or smoke. While radar and sonar provide complementary sensing capabilities, their inherent sparsity and noise present fundamental challenges for direct fusion with dense light detection and ranging (LiDAR) representations. Deep learning-based fusion of heterogeneous sensor modalities in a shared embedding space enables seamless translation from sparse measurements to dense LiDAR-like predictions. However, naive frame-independent fusion suffers from temporal inconsistencies, geometric flickering, and unstable features, which catastrophically degrade downstream simultaneous localization and mapping (SLAM) performance. To address this, we introduce a novel approach to temporal coherence through three complementary mechanisms: explicit temporal regularization encouraging smooth latent transitions, motion-aware transformation losses supervising perceptual displacement consistency, and learned temporal fusion that aggregates information across time. We formalize the latent temporal coherence condition as a desired condition and demonstrate systematic violations in baseline models that are eliminated by our method. Evaluation on real-world indoor datasets shows substantial SLAM improvements. We further contribute domain-adapted metrics (Fréschet video motion distance and correlation-peak distance) that correlate strongly with navigation performance. The framework maintains plug-and-play modularity while ensuring the temporal stability required for reliable autonomous systems.
dc.description.wosFundingTextThis work was supported by the Research Foundation Flanders (FWO) under Grant 1S75624N.
dc.identifier.doi10.1109/tro.2026.3710400
dc.identifier.eissn1941-0468
dc.identifier.issn1552-3098
dc.identifier.urihttps://imec-publications.be/handle/20.500.12860/60433
dc.language.isoeng
dc.provenance.editstepusergreet.vanhoof@imec.be
dc.publisherIEEE-INST ELECTRICAL ELECTRONICS ENGINEERS INC
dc.source.beginpage2861
dc.source.endpage2876
dc.source.journalIEEE TRANSACTIONS ON ROBOTICS
dc.source.numberofpages16
dc.source.volume42
dc.title

LiDAR-BIND-T: Temporally Consistent Sensor Modality Translation and Fusion for Robotic Applications

dc.typeJournal article
dspace.entity.typePublication
imec.internal.crawledAt2026-07-07
imec.internal.sourcecrawler
imec.internal.wosCreatedAt2026-09-07
Files
Publication available in collections: