Qualitative Results
TL;DR
SANDesc can be trained on top of any keypoint detector, leaving the detector untouched. It matches state-of-the-art descriptors on low-resolution data and surpasses all of them on high-resolution data, while still running on a single consumer GPU with 24 GB of VRAM.
Contributions
- We propose SANDesc, a Streamlined Attention-based Network for Descriptor extraction that can be trained on top of any existing keypoint detector, improving matching without modifying the detector itself.
- We design a revised U-Net-like architecture built from Residual U-Net Blocks with Attention (Convolutional Block Attention Modules plus residual paths), reaching strong local representation with just 2.4M parameters.
- We train with a modified triplet loss combined with a curriculum-learning–inspired hard negative mining strategy, which improves training stability.
- We show on HPatches, MegaDepth-1500, and the Image Matching Challenge 2021 that SANDesc is on par with state-of-the-art descriptors at low resolution, and we introduce a new urban 4K dataset with pre-calibrated intrinsics on which it substantially outperforms them — all while training and running on a single 24 GB consumer GPU.
Quantitative Results
Graz4K @ 4K
MegaDepth-1500 @ 2048 kpts
IMC 2021 (avg over 9 scenes)
SuperPoint
DISK
RIPE
ALIKED
DeDoDe-G
+SANDesc
AUC@5 (higher is better) across three matching benchmarks. DeDoDe-G runs out of memory on Graz4K at 4K.
Graz4K Sparse Reconstructions
BibTeX
@inproceedings{durso2026sandesc,
title={A Streamlined Attention-based Network for Descriptor Extraction},
author={D'Urso, Mattia and Santellani, Emanuele and Sormann, Christian and Rossi, Mattia and Kuhn, Andreas and Fraundorfer, Friedrich},
booktitle={2026 International Conference on 3D Vision (3DV)},
year={2026},
organization={IEEE Computer Society}
}