liu.seSearch for publications in DiVA
Change search
Link to record
Permanent link

Direct link
Felsberg, Michael, ProfessorORCID iD iconorcid.org/0000-0002-6096-3648
Alternative names
Publications (10 of 226) Show all publications
Athanasiadis, I., Karmush, A. & Felsberg, M. (2026). Grounding Functional Similarity by Invariance-Aware Model Stitching. In: Proceedings of the 43rd International Conference on Machine Learning: . Paper presented at Proceedings of the 43rd International Conference on Machine Learning, Seoul, South Korea, July 6th - 11th, 2026 (pp. 1-22). San Diego: The International Conference on Machine Learning (ICML)
Open this publication in new window or tab >>Grounding Functional Similarity by Invariance-Aware Model Stitching
2026 (English)In: Proceedings of the 43rd International Conference on Machine Learning, San Diego: The International Conference on Machine Learning (ICML) , 2026, p. 1-22Conference paper, Published paper (Refereed)
Abstract [en]

In deep learning, functional similarity evaluation quantifies the extent to which independently trained models learn similar input--output relationships. In model stitching, functional similarity is framed as representation forward compatibility, i.e., whether the representations of two models can be aligned to solve a given task. Recent studies, however, highlight a critical limitation: models relying on different information cues can still produce compatible representations, making them appear misleadingly similar (Smith et al., 2025). We attribute this failure to standard model stitching being inherently blind to the invariance properties of the stitched models. To address this limitation, we introduce the forward--backward compatibility requirement under which we formulate the invariance-aware model stitching. Through analyzing key stitching configurations, we study the interplay between forward and backward compatibility, showing that invariance-aware model stitching provides a more principled approach to functional similarity evaluation while revealing functional discrepancies previously obscured.

Place, publisher, year, edition, pages
San Diego: The International Conference on Machine Learning (ICML), 2026
Keywords
model stitching, functional similarity evaluation, representation learning
National Category
Computer Vision and Learning Systems
Identifiers
urn:nbn:se:liu:diva-226318 (URN)
Conference
Proceedings of the 43rd International Conference on Machine Learning, Seoul, South Korea, July 6th - 11th, 2026
Funder
Wallenberg AI, Autonomous Systems and Software Program (WASP)
Note

The full text is published under the CC BY 4.0 license.

https://creativecommons.org/licenses/by/4.0/

Available from: 2026-08-03 Created: 2026-08-03 Last updated: 2026-08-03Bibliographically approved
Shafique, B. S., Vayani, A., Maaz, M., Rasheed, H. A., Dissanayake, D., Kurpath, M. I., . . . Khan, F. S. (2025). A Culturally-diverse Multilingual Multimodal Video Benchmark & Model. In: Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng (Ed.), Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: . Paper presented at 2025 Conference on Empirical Methods in Natural Language Processing, Suzhou, China, November 4-9, 2025 (pp. 19998-20022). Association for Computational Linguistics
Open this publication in new window or tab >>A Culturally-diverse Multilingual Multimodal Video Benchmark & Model
Show others...
2025 (English)In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing / [ed] Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng, Association for Computational Linguistics , 2025, p. 19998-20022Conference paper, Published paper (Other academic)
Abstract [en]

Large multimodal models (LMMs) have recently gained attention due to their effectiveness to understand and generate descriptions of visual content. Most existing LMMs are in English language. While few recent works explore multilingual image LMMs, to the best of our knowledge, moving beyond the English language for cultural and linguistic inclusivity is yet to be investigated in the context of video LMMs. In pursuit of more inclusive video LMMs, we introduce a multilingual Video LMM benchmark, named ViMUL-Bench, to evaluate Video LMMs across 14 languages, including both low- and high-resource languages: Arabic, Bengali, Chinese, English, French, German, Hindi, Japanese, Russian, Sinhala, Spanish, Swedish, Tamil, and Urdu. Our ViMUL-Bench is designed to rigorously test video LMMs across 15 categories including eight culturally diverse categories, ranging from lifestyles and festivals to foods and rituals and from local landmarks to prominent cultural personalities. ViMUL-Bench comprises both open-ended (short and long-form) and multiple-choice questions spanning various video durations (short, medium, and long) with 8k samples that are manually verified by native language speakers. In addition, we also introduce a machine translated multilingual video training set comprising 1.2 million samples and develop a simple multilingual video LMM, named ViMUL, that is shown to provide a better tradeoff between high-and low-resource languages for video understanding. We hope our ViMUL-Bench and multilingual video LMM along with a large-scale multilingual video training set will help ease future research in developing cultural and linguistic inclusive multilingual video LMMs. Our proposed benchmark, video LMM and training data will be publicly released.

Place, publisher, year, edition, pages
Association for Computational Linguistics, 2025
National Category
Other Engineering and Technologies
Identifiers
urn:nbn:se:liu:diva-223768 (URN)10.18653/v1/2025.emnlp-main.1012 (DOI)
Conference
2025 Conference on Empirical Methods in Natural Language Processing, Suzhou, China, November 4-9, 2025
Available from: 2026-05-11 Created: 2026-05-11 Last updated: 2026-05-11
Athanasiadis, I., Lindsten, F. & Felsberg, M. (2025). Prior Learning in Introspective VAEs. Transactions on Machine Learning Research (06), 1-41
Open this publication in new window or tab >>Prior Learning in Introspective VAEs
2025 (English)In: Transactions on Machine Learning Research, E-ISSN 2835-8856, no 06, p. 1-41Article in journal (Refereed) Published
Abstract [en]

Variational Autoencoders (VAEs) are a popular framework for unsupervised learning and data generation. A plethora of methods have been proposed focusing on improving VAEs,with the incorporation of adversarial objectives and the integration of prior learning mechanismsbeing prominent directions. When it comes to the former, an indicative instance is therecently introduced family of Introspective VAEs aiming at ensuring that a low likelihood isassigned to unrealistic samples. In this study, we focus on the Soft-IntroVAE (S-IntroVAE),one of only two members of the Introspective VAE family, the other being the originalIntroVAE. We select S-IntroVAE for its state-of-the-art status and its training stability.In particular, we investigate the implication of incorporating a multimodal and trainableprior into this S-IntroVAE. Namely, we formulate the prior as a third player and show thatwhen trained in cooperation with the decoder constitutes an effective way for prior learning,which shares the Nash Equilibrium with the vanilla S-IntroVAE. Furthermore, basedon a modified formulation of the optimal ELBO in S-IntroVAE, we develop theoreticallymotivated regularizations, namely (i) adaptive variance clipping to stabilize training whenlearning the prior and (ii) responsibility regularization to discourage the formation of inactiveprior modes. Finally, we perform a series of targeted experiments on a 2D densityestimation benchmark and in an image generation setting comprised of the (F)-MNIST andCIFAR-10 datasets demonstrating the effect of prior learning in S-IntroVAE in generationand representation learning.

Keywords
probaiblity theory and statistics, computer sciences
National Category
Computer Sciences Computer Vision and Learning Systems
Identifiers
urn:nbn:se:liu:diva-214678 (URN)2-s2.0-105007990039 (Scopus ID)
Funder
Wallenberg AI, Autonomous Systems and Software Program (WASP)
Available from: 2025-06-12 Created: 2025-06-12 Last updated: 2026-06-26Bibliographically approved
Jonnarth, A., Johansson, O., Zhao, J. & Felsberg, M. (2025). Sim-to-Real Transfer of Deep Reinforcement Learning Agents for Online Coverage Path Planning. IEEE Access, 13, 106883-106905
Open this publication in new window or tab >>Sim-to-Real Transfer of Deep Reinforcement Learning Agents for Online Coverage Path Planning
2025 (English)In: IEEE Access, E-ISSN 2169-3536, Vol. 13, p. 106883-106905Article in journal (Refereed) Published
Abstract [en]

Coverage path planning (CPP) is the problem of finding a path that covers the entire free space of a confined area, with applications ranging from robotic lawn mowing to search-and-rescue. While for known environments, offline methods can find provably complete paths, and in some cases optimal solutions, unknown environments need to be planned online during mapping. We investigate the suitability of continuous-space reinforcement learning (RL) for this challenging problem, and propose a computationally feasible egocentric map representation based on frontiers, as well as a novel reward term based on total variation to promote complete coverage. Compared to existing classical methods, this approach allows for a flexible path space, and enables the agent to adapt to specific environment characteristics. Meanwhile, the deployment of RL models on real robot systems is difficult. Training from scratch may be infeasible due to slow convergence times, while transferring from simulation to reality, i.e. sim-to-real transfer, is a key challenge in itself. We bridge the sim-to-real gap through a semi-virtual environment, including a real robot and real-time aspects, while utilizing a simulated sensor and obstacles to enable environment randomization and automated episode resetting. We investigate what level of fine-tuning is needed for adapting to a realistic setting. Through extensive experiments, we show that our approach surpasses the performance of both previous RL-based approaches and highly specialized methods across multiple CPP variations in simulation. Meanwhile, our method successfully transfers to a real robot. Our code implementation can be found online (Link to code repository: https://github.com/arvijj/rl-cpp).

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2025
National Category
Robotics and automation
Identifiers
urn:nbn:se:liu:diva-216664 (URN)10.1109/access.2025.3581035 (DOI)001516410000036 ()2-s2.0-105008639074 (Scopus ID)
Funder
Vinnova, Dnr 2022-02678
Note

Funding Agencies|Wallenberg AI, Autonomous Systems and Software Program (WASP); Knut and Alice Wallenberg (KAW) Foundation; Vinnova Project; Human-Centered Autonomous Regional Airport [2022-02678]; Strategic Research Environment, Excellence Center at Linkoeping-Lund in Information Technology (ELLIIT); Swedish Government; Vinnova [2022-02678] Funding Source: Vinnova

Available from: 2025-08-21 Created: 2025-08-21 Last updated: 2025-08-29
Kristan, M., Matas, J., Tokmakov, P., Lukežič, A., Felsberg, M., Zajc, L. Č., . . . Zhu, X. (2025). The Third Visual Object Tracking Segmentation VOTS2025 Challenge Results. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops: . Paper presented at 2025 International Conference on Computer Vision Workshops-ICCVW, Honolulu, HI, oct 19-20, 2025 (pp. 7481-7499). Computer Vision Foundation
Open this publication in new window or tab >>The Third Visual Object Tracking Segmentation VOTS2025 Challenge Results
Show others...
2025 (English)In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops, Computer Vision Foundation , 2025, p. 7481-7499Conference paper, Published paper (Refereed)
Abstract [en]

The VOTS2025 challenge marks the thirteenth edition of the Visual Object Tracking Segmentation benchmarking activity organized under the VOT initiative. Building on the tracking setup introduced in VOTS2023, the challenge continues to integrate short-term and long-term tracking, as well as single-target and multi-target scenarios, using segmentation masks as the sole form of target annotation. This year's benchmark features three sub-challenges. The first two, VOTS2025 and VOTSt2025, evaluate tracking of conventional objects and objects undergoing topological changes, respectively. A new addition, VOTS-RT2025, aims to foster the development of efficient tracking models by introducing constraints that highlight realtime performance. All sub-challenges adopt a consistent evaluation protocol, with VOTS-RT2025 introducing specific modifications to reflect latency-aware performance. We report and analyze results from 32 submissions. Full tracker descriptions, source code, datasets, and the evaluation toolkit are available on the project website https://www.votchallenge.net/vots2025.

Place, publisher, year, edition, pages
Computer Vision Foundation, 2025
Series
IEEE International Conference on Computer Vision Workshops, ISSN 2473-9936, E-ISSN 2473-9944
National Category
Computer Vision and Learning Systems
Identifiers
urn:nbn:se:liu:diva-223771 (URN)10.1109/ICCVW69036.2025.00771 (DOI)001740020100004 ()2-s2.0-105035209983 (Scopus ID)9798331589882 (ISBN)9798331589899 (ISBN)
Conference
2025 International Conference on Computer Vision Workshops-ICCVW, Honolulu, HI, oct 19-20, 2025
Available from: 2026-05-11 Created: 2026-05-11 Last updated: 2026-05-26
Dehdarirad, T., Eilertsen, G. & Felsberg, M. (2025). When Non-Commutativity Breeds Unfairness: A Geometric–Algebraic View of Uncertainty in VAEs. In: EurIPS 2025 Workshop -- Unifying Perspectives on Learning Biases: . Paper presented at EurIPS 2025 Workshop -- Unifying Perspectives on Learning Biases.
Open this publication in new window or tab >>When Non-Commutativity Breeds Unfairness: A Geometric–Algebraic View of Uncertainty in VAEs
2025 (English)In: EurIPS 2025 Workshop -- Unifying Perspectives on Learning Biases, 2025Conference paper, Poster (with or without abstract) (Other academic)
Abstract [en]

We propose a novel theoretical framework that unifies geometric and algebraic perspectives on uncertainty quantification in deep generative models. Standard variational autoencoders (VAEs) often underestimate uncertainty in the presence of non-commutative transformations, leading to miscalibrated confidence and potential fairness violations. Our approach introduces an integrated diagnostic and regularization framework that monitors transformation relationships via Baker--Campbell--Hausdorff deviation and decoder order-swap tests; then, we apply category-specific regularization. Commutative pairs receive standard penalties, decoder-induced artifacts are suppressed through equivariance constraints, and genuinely non-commutative transformations are governed by a deformation-stability principle linking commutator strength to required uncertainty scaling. The key theoretical result establishes that geometric uncertainty, measured through the Riemannian sensitivity of the decoder to Lie-algebra actions, should scale with algebraic non-commutativity to ensure proper calibration. This framework closes a gap in generative modelling by providing principled diagnostics and regularization strategies for geometrically-aware uncertainty in symmetry-rich latent spaces, with direct implications for fairness when biased correlations induce spurious non-commutativity.

National Category
Algorithms Algebra and Logic Computer Vision and Learning Systems
Identifiers
urn:nbn:se:liu:diva-223850 (URN)
Conference
EurIPS 2025 Workshop -- Unifying Perspectives on Learning Biases
Available from: 2026-05-11 Created: 2026-05-11 Last updated: 2026-05-29
Edstedt, J., Bökman, G., Wadenbäck, M. & Felsberg, M. (2024). DeDoDe: Detect, Don't Describe — Describe, Don't Detect for Local Feature Matching. In: 2024 International Conference on 3D Vision (3DV): . Paper presented at International Conference on 3D Vision (3DV), Davos, Switzerland, 18-21 March, 2024.. Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>DeDoDe: Detect, Don't Describe — Describe, Don't Detect for Local Feature Matching
2024 (English)In: 2024 International Conference on 3D Vision (3DV), Institute of Electrical and Electronics Engineers (IEEE), 2024Conference paper, Published paper (Refereed)
Abstract [en]

Keypoint detection is a pivotal step in 3D reconstruction, whereby sets of (up to) K points are detected in each view of a scene. Crucially, the detected points need to be consistent between views, i.e., correspond to the same 3D point in the scene. One of the main challenges with keypoint detection is the formulation of the learning objective. Previous learning-based methods typically jointly learn descriptors with keypoints, and treat the keypoint detection as a binary classification task on mutual nearest neighbours. However, basing keypoint detection on descriptor nearest neighbours is a proxy task, which is not guaranteed to produce 3D-consistent keypoints. Furthermore, this ties the keypoints to a specific descriptor, complicating downstream usage. In this work, we instead learn keypoints directly from 3D consistency. To this end, we train the detector to detect tracks from large-scale SfM. As these points are often overly sparse, we derive a semi-supervised two-view detection objective to expand this set to a desired number of detections. To train a descriptor, we maximize the mutual nearest neighbour objective over the keypoints with a separate network. Results show that our approach, DeDoDe, achieves significant gains on multiple geometry benchmarks. Code is provided at http://github.com/Parskatt/DeDoDegithub.com/Parskatt/DeDoDe.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2024
Series
2024 International Conference on 3D Vision (3DV), ISSN 2378-3826, E-ISSN 2475-7888
National Category
Computer Vision and Learning Systems Engineering and Technology
Identifiers
urn:nbn:se:liu:diva-204892 (URN)10.1109/3dv62453.2024.00035 (DOI)001250581700028 ()2-s2.0-85180219084 (Scopus ID)9798350362459 (ISBN)9798350362466 (ISBN)
Conference
International Conference on 3D Vision (3DV), Davos, Switzerland, 18-21 March, 2024.
Note

Funding Agencies|Wallenberg Artificial Intelligence, Autonomous Systems and Software Program (WASP) - Knut and Alice Wallenberg Foundation; strategic research environment ELLIIT - Swedish government; Swedish Research Council [2022-06725]; Knut and Alice Wallenberg Foundation at the National Supercomputer Centre

Available from: 2024-06-17 Created: 2024-06-17 Last updated: 2026-06-23
Sanchez Aimar, E., Helgesen, N., Xu, Y., Kuhlmann, M. & Felsberg, M. (2024). Flexible Distribution Alignment: Towards Long-Tailed Semi-supervised Learning with Proper Calibration. In: Aleš Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, Gül Varol (Ed.), Computer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings, Part LIV. Paper presented at 18th European Conference, Milan, Italy, September 29–October 4, 2024 (pp. 307-327). Springer Nature Switzerland, 15112
Open this publication in new window or tab >>Flexible Distribution Alignment: Towards Long-Tailed Semi-supervised Learning with Proper Calibration
Show others...
2024 (English)In: Computer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings, Part LIV / [ed] Aleš Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, Gül Varol, Springer Nature Switzerland , 2024, Vol. 15112, p. 307-327Conference paper, Published paper (Refereed)
Abstract [en]

Long-tailed semi-supervised learning (LTSSL) represents a practical scenario for semi-supervised applications, challenged by skewed labeled distributions that bias classifiers. This problem is often aggravated by discrepancies between labeled and unlabeled class distributions, leading to biased pseudo-labels, neglect of rare classes, and poorly calibrated probabilities. To address these issues, we introduce Flexible Distribution Alignment (FlexDA), a novel adaptive logit-adjusted loss framework designed to dynamically estimate and align predictions with the actual distribution of unlabeled data and achieve a balanced classifier by the end of training. FlexDA is further enhanced by a distillation-based consistency loss, promoting fair data usage across classes and effectively leveraging underconfident samples. This method, encapsulated in ADELLO (Align and Distill Everything All at Once), proves robust against label shift, significantly improves model calibration in LTSSL contexts, and surpasses previous state-of-of-art approaches across multiple benchmarks, including CIFAR100-LT, STL10-LT, and ImageNet127, addressing class imbalance challenges in semi-supervised learning. Our code is available at https://github.com/emasa/ADELLO-LTSSL.

Place, publisher, year, edition, pages
Springer Nature Switzerland, 2024
Series
Lecture Notes in Computer Science, ISSN 0302-9743, E-ISSN 1611-3349 ; 15112
National Category
Computer Vision and Learning Systems
Identifiers
urn:nbn:se:liu:diva-209223 (URN)10.1007/978-3-031-72949-2_18 (DOI)001352860600018 ()2-s2.0-85208545165 (Scopus ID)9783031729485 (ISBN)9783031729492 (ISBN)
Conference
18th European Conference, Milan, Italy, September 29–October 4, 2024
Note

Funding Agencies|Wallenberg Artificial Intelligence, Autonomous Systems and Software Program (WASP) - Knut and Alice Wallenberg Foundation; Swedish Research Council [2022-06725]; Knut and Alice Wallenberg Foundation at the National Supercomputer Centre

Available from: 2024-11-06 Created: 2024-11-06 Last updated: 2026-06-24
Jonnarth, A., Zhang, Y. & Felsberg, M. (2024). High-fidelity Pseudo-labels for Boosting Weakly-Supervised Segmentation. In: 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV): . Paper presented at IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA, jan 3-8, 2024 (pp. 999-1008). Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>High-fidelity Pseudo-labels for Boosting Weakly-Supervised Segmentation
2024 (English)In: 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Institute of Electrical and Electronics Engineers (IEEE), 2024, p. 999-1008Conference paper, Published paper (Refereed)
Abstract [en]

Image-level weakly-supervised semantic segmentation (WSSS) reduces the usually vast data annotation cost by surrogate segmentation masks during training. The typical approach involves training an image classification network using global average pooling (GAP) on convolutional feature maps. This enables the estimation of object locations based on class activation maps (CAMs), which identify the importance of image regions. The CAMs are then used to generate pseudo-labels, in the form of segmentation masks, to supervise a segmentation model in the absence of pixel-level ground truth. Our work is based on two techniques for improving CAMs; importance sampling, which is a substitute for GAP, and the feature similarity loss, which utilizes a heuristic that object contours almost always align with color edges in images. However, both are based on the multinomial posterior with softmax, and implicitly assume that classes are mutually exclusive, which turns out suboptimal in our experiments. Thus, we reformulate both techniques based on binomial posteriors of multiple independent binary problems. This has two benefits; their performance is improved and they become more general, resulting in an add-on method that can boost virtually any WSSS method. This is demonstrated on a wide variety of baselines on the PASCAL VOC dataset, improving the region similarity and contour quality of all implemented state-of-the-art methods. Experiments on the MS COCO dataset further show that our proposed add-on is well-suited for large-scale settings. Our code implementation is available at https://github.com/arvijj/hfpl.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2024
Keywords
weakly supervised, semantic segmentation, importance sampling, feature similarity, class activation maps
National Category
Computer Vision and Learning Systems
Identifiers
urn:nbn:se:liu:diva-202446 (URN)10.1109/WACV57701.2024.00105 (DOI)001222964601011 ()2-s2.0-85191946457 (Scopus ID)
Conference
IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA, jan 3-8, 2024
Note

Funding Agencies|Wallenberg AI, Autonomous Systems and Software Program (WASP) - Knut and Alice Wallenberg (KAW) Foundation; Swedish Research Council [2022-06725]

Available from: 2024-04-15 Created: 2024-04-15 Last updated: 2026-06-22Bibliographically approved
Jonnarth, A., Zhao, J. & Felsberg, M. (2024). Learning Coverage Paths in Unknown Environments with Deep Reinforcement Learning. In: Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, Felix Berkenkamp (Ed.), Proceedings of the 41st International Conference on Machine Learning: . Paper presented at International Conference on Machine Learning, 21-27 July 2024, Vienna, Austria (pp. 22491-22508). PMLR
Open this publication in new window or tab >>Learning Coverage Paths in Unknown Environments with Deep Reinforcement Learning
2024 (English)In: Proceedings of the 41st International Conference on Machine Learning / [ed] Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, Felix Berkenkamp, PMLR , 2024, p. 22491-22508Conference paper, Published paper (Refereed)
Abstract [en]

Coverage path planning (CPP) is the problem of finding a path that covers the entire free space of a confined area, with applications ranging from robotic lawn mowing to search-and-rescue. When the environment is unknown, the path needs to be planned online while mapping the environment, which cannot be addressed by offline planning methods that do not allow for a flexible path space. We investigate how suitable reinforcement learning is for this challenging problem, and analyze the involved components required to efficiently learn coverage paths, such as action space, input feature representation, neural network architecture, and reward function. We propose a computationally feasible egocentric map representation based on frontiers, and a novel reward term based on total variation to promote complete coverage. Through extensive experiments, we show that our approach surpasses the performance of both previous RL-based approaches and highly specialized methods across multiple CPP variations.

Place, publisher, year, edition, pages
PMLR, 2024
Series
Proceedings of Machine Learning Research, ISSN 2640-3498 ; 235
National Category
Computer Sciences Computer Vision and Learning Systems
Identifiers
urn:nbn:se:liu:diva-207087 (URN)
Conference
International Conference on Machine Learning, 21-27 July 2024, Vienna, Austria
Note

Funding agencies: y the Wallenberg AI, Autonomous Systems and Software Program (WASP), fundedby the Knut and Alice Wallenberg (KAW) Foundation;  the Vinnova project, human centered autonomous regional airport, Dnr 2022-02678. The computational resources were provided by the National Academic Infrastructure for Supercomputing in Sweden (NAISS), partially funded by the Swedish Research Council through grant agreement no. 2022-06725, and by the Berzelius resource, provided by the KAW Foundation at the National Supercomputer Centre (NSC). 

Available from: 2024-08-30 Created: 2024-08-30 Last updated: 2026-06-22
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0002-6096-3648

Search in DiVA

Show all publications