liu.seSearch for publications in DiVA
Change search
Link to record
Permanent link

Direct link
Publications (10 of 10) Show all publications
Edstedt, J., Bökman, G., Wadenbäck, M. & Felsberg, M. (2024). DeDoDe: Detect, Don't Describe — Describe, Don't Detect for Local Feature Matching. In: 2024 International Conference on 3D Vision (3DV): . Paper presented at International Conference on 3D Vision (3DV), Davos, Switzerland, 18-21 March, 2024.. Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>DeDoDe: Detect, Don't Describe — Describe, Don't Detect for Local Feature Matching
2024 (English)In: 2024 International Conference on 3D Vision (3DV), Institute of Electrical and Electronics Engineers (IEEE), 2024Conference paper, Published paper (Refereed)
Abstract [en]

Keypoint detection is a pivotal step in 3D reconstruction, whereby sets of (up to) K points are detected in each view of a scene. Crucially, the detected points need to be consistent between views, i.e., correspond to the same 3D point in the scene. One of the main challenges with keypoint detection is the formulation of the learning objective. Previous learning-based methods typically jointly learn descriptors with keypoints, and treat the keypoint detection as a binary classification task on mutual nearest neighbours. However, basing keypoint detection on descriptor nearest neighbours is a proxy task, which is not guaranteed to produce 3D-consistent keypoints. Furthermore, this ties the keypoints to a specific descriptor, complicating downstream usage. In this work, we instead learn keypoints directly from 3D consistency. To this end, we train the detector to detect tracks from large-scale SfM. As these points are often overly sparse, we derive a semi-supervised two-view detection objective to expand this set to a desired number of detections. To train a descriptor, we maximize the mutual nearest neighbour objective over the keypoints with a separate network. Results show that our approach, DeDoDe, achieves significant gains on multiple geometry benchmarks. Code is provided at http://github.com/Parskatt/DeDoDegithub.com/Parskatt/DeDoDe.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2024
Series
2024 International Conference on 3D Vision (3DV), ISSN 2378-3826, E-ISSN 2475-7888
National Category
Computer Vision and Learning Systems Engineering and Technology
Identifiers
urn:nbn:se:liu:diva-204892 (URN)10.1109/3dv62453.2024.00035 (DOI)001250581700028 ()2-s2.0-85180219084 (Scopus ID)9798350362459 (ISBN)9798350362466 (ISBN)
Conference
International Conference on 3D Vision (3DV), Davos, Switzerland, 18-21 March, 2024.
Note

Funding Agencies|Wallenberg Artificial Intelligence, Autonomous Systems and Software Program (WASP) - Knut and Alice Wallenberg Foundation; strategic research environment ELLIIT - Swedish government; Swedish Research Council [2022-06725]; Knut and Alice Wallenberg Foundation at the National Supercomputer Centre

Available from: 2024-06-17 Created: 2024-06-17 Last updated: 2026-06-23
Melnyk, P., Felsberg, M., Wadenbäck, M., Robinson, A. & Le, C. (2024). On Learning Deep O(n)-Equivariant Hyperspheres. In: Salakhutdinov, Ruslan and Kolter, Zico and Heller, Katherine and Weller, Adrian and Oliver, Nuria and Scarlett, Jonathan and Berkenkamp, Felix (Ed.), Proceedings of the 41st International Conference on Machine Learning: . Paper presented at 41st International Conference on Machine Learning, Vienna, Austria, 21-27 July 2024 (pp. 35324-35339). PMLR, 235
Open this publication in new window or tab >>On Learning Deep O(n)-Equivariant Hyperspheres
Show others...
2024 (English)In: Proceedings of the 41st International Conference on Machine Learning / [ed] Salakhutdinov, Ruslan and Kolter, Zico and Heller, Katherine and Weller, Adrian and Oliver, Nuria and Scarlett, Jonathan and Berkenkamp, Felix, PMLR , 2024, Vol. 235, p. 35324-35339Conference paper, Poster (with or without abstract) (Refereed)
Abstract [en]

In this paper, we utilize hyperspheres and regular n-simplexes and propose an approach to learning deep features equivariant under the transformations of nD reflections and rotations, encompassed by the powerful group of O(n). Namely, we propose O(n)-equivariant neurons with spherical decision surfaces that generalize to any dimension n, which we call Deep Equivariant Hyperspheres. We demonstrate how to combine them in a network that directly operates on the basis of the input points and propose an invariant operator based on the relation between two points and a sphere, which as we show, turns out to be a Gram matrix. Using synthetic and real-world data in nD, we experimentally verify our theoretical contributions and find that our approach is superior to the competing methods for O(n)-equivariant benchmark datasets (classification and regression), demonstrating a favorable speed/performance trade-off. The code is available on GitHub.

Place, publisher, year, edition, pages
PMLR, 2024
Series
Proceedings of Machine Learning Research, ISSN 2640-3498 ; 235
National Category
Computer Vision and Learning Systems Engineering and Technology
Identifiers
urn:nbn:se:liu:diva-206460 (URN)
Conference
41st International Conference on Machine Learning, Vienna, Austria, 21-27 July 2024
Available from: 2024-08-14 Created: 2024-08-14 Last updated: 2026-04-29
Edstedt, J., Sun, Q., Bökman, G., Wadenbäck, M. & Felsberg, M. (2024). RoMa: Robust Dense Feature Matching. In: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR): . Paper presented at 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Seattle, WA, USA, 16-22 June 2024. (pp. 19790-19800). Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>RoMa: Robust Dense Feature Matching
Show others...
2024 (English)In: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Institute of Electrical and Electronics Engineers (IEEE), 2024, p. 19790-19800Conference paper, Published paper (Refereed)
Abstract [en]

Feature matching is an important computer vision task that involves estimating correspondences between two images of a 3D scene, and dense methods estimate all such correspondences. The aim is to learn a robust model, i.e., a model able to match under challenging real-world changes. In this work, we propose such a model, leveraging frozen pretrained features from the foundation model DINOv2. Al-though these features are significantly more robust than local features trained from scratch, they are inherently coarse. We therefore combine them with specialized ConvNet fine features, creating a precisely localizable feature pyramid. To further improve robustness, we propose a tailored transformer match decoder that predicts anchor probabilities, which enables it to express multimodality. Finally, we propose an improved loss formulation through regression-by-classification with subsequent robust regression. We conduct a comprehensive set of experiments that show that our method, RoMa, achieves significant gains, setting a new state-of-the-art. In particular, we achieve a 36% improvement on the extremely challenging WxBS benchmark. Code is provided at github.com/Parskatt/RoMa.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2024
Series
Conference on Computer Vision and Pattern Recognition (CVPR), ISSN 1063-6919, E-ISSN 2575-7075
National Category
Computer Vision and Learning Systems Engineering and Technology
Identifiers
urn:nbn:se:liu:diva-207702 (URN)10.1109/CVPR52733.2024.01871 (DOI)001342515503014 ()2-s2.0-85199525100 (Scopus ID)9798350353006 (ISBN)9798350353013 (ISBN)
Conference
2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Seattle, WA, USA, 16-22 June 2024.
Note

Funding Agencies|Wallenberg Artificial Intelligence, Autonomous Systems and Software Program (WASP) - Knut and Alice Wallenberg Foundation; strategic research environment ELLIIT - Swedish government; Swedish Research Council [2022-06725]; Knut and Alice Wallenberg Foundation at the National Supercomputer Centre

Available from: 2024-09-17 Created: 2024-09-17 Last updated: 2026-04-29
Melnyk, P., Robinson, A., Felsberg, M. & Wadenbäck, M. (2024). TetraSphere: A Neural Descriptor for O(3)-Invariant Point Cloud Analysis. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). June 2024: . Paper presented at IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16-22 June, 2024. (pp. 5620-5630). IEEE Computer Society
Open this publication in new window or tab >>TetraSphere: A Neural Descriptor for O(3)-Invariant Point Cloud Analysis
2024 (English)In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). June 2024, IEEE Computer Society, 2024, p. 5620-5630Conference paper, Published paper (Refereed)
Abstract [en]

In many practical applications, 3D point cloud analysis requires rotation invariance. In this paper, we present a learnable descriptor invariant under 3D rotations and reflections, i.e., the O(3) actions, utilizing the recently introduced steerable 3D spherical neurons and vector neurons. Specifically, we propose an embedding of the 3D spherical neurons into 4D vector neurons, which leverages end-to-end training of the model. In our approach, we perform TetraTransform--an equivariant embedding of the 3D input into 4D, constructed from the steerable neurons--and extract deeper O(3)-equivariant features using vector neurons. This integration of the TetraTransform into the VN-DGCNN framework, termed TetraSphere, negligibly increases the number of parameters by less than 0.0002%. TetraSphere sets a new state-of-the-art performance classifying randomly rotated real-world object scans of the challenging subsets of ScanObjectNN. Additionally, TetraSphere outperforms all equivariant methods on randomly rotated synthetic data: classifying objects from ModelNet40 and segmenting parts of the ShapeNet shapes. Thus, our results reveal the practical value of steerable 3D spherical neurons for learning in 3D Euclidean space

Place, publisher, year, edition, pages
IEEE Computer Society, 2024
Series
IEEE Conference on Computer Vision and Pattern Recognition, ISSN 1063-6919, E-ISSN 2575-7075
National Category
Computer Vision and Learning Systems Engineering and Technology
Identifiers
urn:nbn:se:liu:diva-207318 (URN)10.1109/CVPR52733.2024.00537 (DOI)001322555906003 ()2-s2.0-85217079523 (Scopus ID)9798350353006 (ISBN)9798350353013 (ISBN)
Conference
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16-22 June, 2024.
Note

Funding Agencies|Wallenberg AI, Autonomous Systems and Software Program (WASP); Swedish Research Council [2022-04266]; strategic research environment EL-LIIT

Available from: 2024-09-04 Created: 2024-09-04 Last updated: 2026-04-29
Edstedt, J., Athanasiadis, I., Wadenbäck, M. & Felsberg, M. (2023). DKM: Dense Kernelized Feature Matching for Geometry Estimation. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR): . Paper presented at 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 17-24 June 2023 (pp. 17765-17775). IEEE Communications Society
Open this publication in new window or tab >>DKM: Dense Kernelized Feature Matching for Geometry Estimation
2023 (English)In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE Communications Society, 2023, p. 17765-17775Conference paper, Published paper (Refereed)
Abstract [en]

Feature matching is a challenging computer vision task that involves finding correspondences between two images of a 3D scene. In this paper we consider the dense approach instead of the more common sparse paradigm, thus striving to find all correspondences. Perhaps counter-intuitively, dense methods have previously shown inferior performance to their sparse and semi-sparse counterparts for estimation of two-view geometry. This changes with our novel dense method, which outperforms both dense and sparse methods on geometry estimation. The novelty is threefold: First, we propose a kernel regression global matcher. Secondly, we propose warp refinement through stacked feature maps and depthwise convolution kernels. Thirdly, we propose learning dense confidence through consistent depth and a balanced sampling approach for dense confidence maps. Through extensive experiments we confirm that our proposed dense method, Dense Kernelized Feature Matching, sets a new state-of-the-art on multiple geometry estimation benchmarks. In particular, we achieve an improvement on MegaDepth-1500 of +4.9 and +8.9 AUC@5° compared to the best previous sparse method and dense method respectively. Our code is provided at the following repository: https://github.com/Parskatt/DKM.

Place, publisher, year, edition, pages
IEEE Communications Society, 2023
Series
Proceedings:IEEE Conference on Computer Vision and Pattern Recognition, ISSN 1063-6919, E-ISSN 2575-7075
National Category
Computer Vision and Learning Systems Engineering and Technology
Identifiers
urn:nbn:se:liu:diva-197717 (URN)10.1109/cvpr52729.2023.01704 (DOI)001062531302008 ()2-s2.0-85215442193 (Scopus ID)9798350301298 (ISBN)9798350301304 (ISBN)
Conference
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 17-24 June 2023
Note

This work was supported by the Wallenberg Artificial Intelligence, Autonomous Systems and Software Program (WASP), funded by Knut and Alice Wallenberg Foundation; andby the strategic research environment ELLIIT funded by the Swedish government. The computational resources were provided by the National Academic Infrastructure forSupercomputing in Sweden (NAISS), partially funded by the Swedish Research Council through grant agreement no. 2022-06725, and by the Berzelius resource, provided bythe Knut and Alice Wallenberg Foundation at the National Supercomputer Centre.

Available from: 2023-09-11 Created: 2023-09-11 Last updated: 2026-04-29Bibliographically approved
Melnyk, P., Felsberg, M. & Wadenbäck, M. (2022). Steerable 3D Spherical Neurons. In: Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, Sivan Sabato (Ed.), Proceedings of the 39th International Conference on Machine Learning: . Paper presented at International Conference on Machine Learning, Baltimore, Maryland, USA, 17-23 July 2022 (pp. 15330-15339). PMLR, 162
Open this publication in new window or tab >>Steerable 3D Spherical Neurons
2022 (English)In: Proceedings of the 39th International Conference on Machine Learning / [ed] Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, Sivan Sabato, PMLR , 2022, Vol. 162, p. 15330-15339Conference paper, Published paper (Refereed)
Abstract [en]

Emerging from low-level vision theory, steerable filters found their counterpart in prior work on steerable convolutional neural networks equivariant to rigid transformations. In our work, we propose a steerable feed-forward learning-based approach that consists of neurons with spherical decision surfaces and operates on point clouds. Such spherical neurons are obtained by conformal embedding of Euclidean space and have recently been revisited in the context of learning representations of point sets. Focusing on 3D geometry, we exploit the isometry property of spherical neurons and derive a 3D steerability constraint. After training spherical neurons to classify point clouds in a canonical orientation, we use a tetrahedron basis to quadruplicate the neurons and construct rotation-equivariant spherical filter banks. We then apply the derived constraint to interpolate the filter bank outputs and, thus, obtain a rotation-invariant network. Finally, we use a synthetic point set and real-world 3D skeleton data to verify our theoretical findings. The code is available at https://github.com/pavlo-melnyk/steerable-3d-neurons.

Place, publisher, year, edition, pages
PMLR, 2022
Series
Proceedings of Machine Learning Research, ISSN 2640-3498
National Category
Computer Vision and Learning Systems Engineering and Technology
Identifiers
urn:nbn:se:liu:diva-187149 (URN)000900064905021 ()
Conference
International Conference on Machine Learning, Baltimore, Maryland, USA, 17-23 July 2022
Note

Funding: Wallenberg AI, Autonomous Systems and Software Program (WASP); Swedish Research Council [2018-04673]; strategic research environment ELLIIT

Available from: 2022-08-08 Created: 2022-08-08 Last updated: 2026-04-29
Valtonen Örnhag, M., Persson, P., Wadenbäck, M., Åström, K. & Heyden, A. (2022). Trust Your IMU: Consequences of Ignoring the IMU Drift. In: Proceedings 2022 IEEE/CVF Conference on Computer Visionand Pattern Recognition Workshops: CVPRW 2022. Paper presented at 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, New Orleans, Louisiana, USA, 19 – 24 June 2022 (pp. 4467-4476). IEEE Computer Society
Open this publication in new window or tab >>Trust Your IMU: Consequences of Ignoring the IMU Drift
Show others...
2022 (English)In: Proceedings 2022 IEEE/CVF Conference on Computer Visionand Pattern Recognition Workshops: CVPRW 2022, IEEE Computer Society, 2022, p. 4467-4476Conference paper, Published paper (Refereed)
Abstract [en]

In this paper, we argue that modern pre-integration methods for inertial measurement units (IMUs) are accurate enough to ignore the drift for short time intervals. This allows us to consider a simplified camera model, which in turn admits further intrinsic calibration. We develop the first-ever solver to jointly solve the relative pose problem with unknown and equal focal length and radial distortion profile while utilizing the IMU data. Furthermore, we show significant speed-up compared to state-of-the-art algorithms, with small or negligible loss in accuracy for partially calibrated setups.The proposed algorithms are tested on both synthetic and real data, where the latter is focused on navigation using unmanned aerial vehicles (UAVs). We evaluate the proposed solvers on different commercially available low-cost UAVs, and demonstrate that the novel assumption on IMU drift is feasible in real-life applications. The extended intrinsic auto-calibration enables us to use distorted input images, making tedious calibration processes obsolete, compared to current state-of-the-art methods. Code available at: https://github.com/marcusvaltonen/DronePoseLib.

Place, publisher, year, edition, pages
IEEE Computer Society, 2022
Series
IEEE Computer Society Conference on Computer Vision and Pattern Recognition workshops, ISSN 2160-7508, E-ISSN 2160-7516
National Category
Computer Vision and Learning Systems Engineering and Technology
Identifiers
urn:nbn:se:liu:diva-188868 (URN)10.1109/cvprw56347.2022.00493 (DOI)000861612704058 ()2-s2.0-85137766516 (Scopus ID)9781665487399 (ISBN)9781665487405 (ISBN)
Conference
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, New Orleans, Louisiana, USA, 19 – 24 June 2022
Note

Funding: Swedish Research Council [2015-05639]; strategic research projects ELLIIT; eSSENCE; Swedish Foundation for Strategic Research project, Semantic Mapping and Visual Navigation for Smart Robots [RIT15-0038]; Wallenberg AI, Autonomous Systems and Software Program (WASP) - Knut and Alice Wallenberg Foundation

Available from: 2022-09-29 Created: 2022-09-29 Last updated: 2026-04-29
Valtonen Örnhag, M., Persson, P., Wadenbäck, M., Åström, K. & Heyden, A. (2021). Efficient Real-Time Radial Distortion Correction for UAVs. In: 2021 IEEE Winter Conference on Applications of Computer Vision (WACV): . Paper presented at 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), 3-8 Jan. 2021 (pp. 1750-1759). IEEE
Open this publication in new window or tab >>Efficient Real-Time Radial Distortion Correction for UAVs
Show others...
2021 (English)In: 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), IEEE , 2021, p. 1750-1759Conference paper, Published paper (Refereed)
Abstract [en]

In this paper we present a novel algorithm for onboard radial distortion correction for unmanned aerial vehicles (UAVs) equipped with an inertial measurement unit (IMU), that runs in real-time. This approach makes calibration procedures redundant, thus allowing for exchange of optics extemporaneously. By utilizing the IMU data, the cameras can be aligned with the gravity direction. This allows us to work with fewer degrees of freedom, and opens up for further intrinsic calibration. We propose a fast and robust minimal solver for simultaneously estimating the focal length, radial distortion profile and motion parameters from homographies. The proposed solver is tested on both synthetic and real data, and perform better or on par with state-of-the-art methods relying on pre-calibration procedures. Code available at: https://github.com/marcusvaltonen/HomLib. 1

Place, publisher, year, edition, pages
IEEE, 2021
Series
IEEE Workshop on Applications of Computer Vision (WACV), ISSN 2472-6737, E-ISSN 2642-9381
National Category
Computer Vision and Learning Systems Engineering and Technology
Identifiers
urn:nbn:se:liu:diva-179283 (URN)10.1109/WACV48630.2021.00179 (DOI)000692171000174 ()2-s2.0-85104186047 (Scopus ID)9781665404778 (ISBN)9781665446402 (ISBN)
Conference
2021 IEEE Winter Conference on Applications of Computer Vision (WACV), 3-8 Jan. 2021
Available from: 2021-09-16 Created: 2021-09-16 Last updated: 2026-04-29
Melnyk, P., Felsberg, M. & Wadenbäck, M. (2021). Embed Me If You Can: A Geometric Perceptron. In: Proceedings 2021 IEEE/CVF International Conference on Computer Vision ICCV 2021: . Paper presented at IEEE/CVF International Conference on Computer Vision (ICCV), 10-17 October 2021 (Virtual Event), Montreal, QC, Canada (pp. 1256-1264). Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>Embed Me If You Can: A Geometric Perceptron
2021 (English)In: Proceedings 2021 IEEE/CVF International Conference on Computer Vision ICCV 2021, Institute of Electrical and Electronics Engineers (IEEE), 2021, p. 1256-1264Conference paper, Published paper (Refereed)
Abstract [en]

Solving geometric tasks involving point clouds by using machine learning is a challenging problem. Standard feed-forward neural networks combine linear or, if the bias parameter is included, affine layers and activation functions. Their geometric modeling is limited, which motivated the prior work introducing the multilayer hypersphere perceptron (MLHP). Its constituent part, i.e., the hypersphere neuron, is obtained by applying a conformal embedding of Euclidean space. By virtue of Clifford algebra, it can be implemented as the Cartesian dot product of inputs and weights. If the embedding is applied in a manner consistent with the dimensionality of the input space geometry, the decision surfaces of the model units become combinations of hyperspheres and make the decision-making process geometrically interpretable for humans. Our extension of the MLHP model, the multilayer geometric perceptron (MLGP), and its respective layer units, i.e., geometric neurons, are consistent with the 3D geometry and provide a geometric handle of the learned coefficients. In particular, the geometric neuron activations are isometric in 3D, which is necessary for rotation and translation equivariance. When classifying the 3D Tetris shapes, we quantitatively show that our model requires no activation function in the hidden layers other than the embedding to outperform the vanilla multilayer perceptron. In the presence of noise in the data, our model is also superior to the MLHP.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2021
Series
IEEE International Conference on Computer Vision. Proceedings, ISSN 1550-5499, E-ISSN 2380-7504
National Category
Computer Vision and Learning Systems Engineering and Technology
Identifiers
urn:nbn:se:liu:diva-183312 (URN)10.1109/iccv48922.2021.00131 (DOI)000797698901044 ()2-s2.0-85119151216 (Scopus ID)9781665428125 (ISBN)9781665428132 (ISBN)
Conference
IEEE/CVF International Conference on Computer Vision (ICCV), 10-17 October 2021 (Virtual Event), Montreal, QC, Canada
Note

Funding: Wallenberg AI, Autonomous Systems and Software Program (WASP); Swedish Research Council [2018-04673]; strategic research environment ELLIIT

Available from: 2022-03-02 Created: 2022-03-02 Last updated: 2026-04-29Bibliographically approved
Edstedt, J., Bökman, G., Wadenbäck, M. & Felsberg, M.DaD: Distilled Reinforcement Learning for Diverse Keypoint Detection.
Open this publication in new window or tab >>DaD: Distilled Reinforcement Learning for Diverse Keypoint Detection
(English)Manuscript (preprint) (Other academic)
Abstract [en]

Keypoints are what enable Structure-from-Motion (SfM) systems to scale to thousands of images. However, designing a keypoint detection objective is a non-trivial task, as SfM is non-differentiable. Typically, an auxiliary objective involving a descriptor is optimized. This however induces a dependency on the descriptor, which is undesirable. In this paper we propose a fully self-supervised and descriptor-free objective for keypoint detection, through reinforcement learning. To ensure training does not degenerate, we leverage a balanced top-K sampling strategy. While this already produces competitive models, we find that two qualitatively different types of detectors emerge, which are only able to detect light and dark keypoints respectively. To remedy this, we train a third detector, DaD, that optimizes the Kullback-Leibler divergence of the pointwise maximum of both light and dark detectors. Our approach significantly improve upon SotA across a range of benchmarks. Code and model weights are publicly available at https://github.com/parskatt/dad

National Category
Computer Vision and Learning Systems Engineering and Technology
Identifiers
urn:nbn:se:liu:diva-217642 (URN)10.48550/arXiv.2503.07347 (DOI)
Available from: 2025-09-11 Created: 2025-09-11 Last updated: 2026-04-29
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0002-0675-2794

Search in DiVA

Show all publications