SANDesc: A Streamlined Attention-based Network for Descriptor Extraction

1Graz University of Technology 2Sony Europe
3DV 2026

Qualitative Results

ALIKED and DeDoDe with SANDesc vs. original descriptors on Brandenburg Gate
ALIKED and DeDoDe with SANDesc descriptors (top) vs. their original descriptors (bottom).
SuperPoint and DISK with SANDesc vs. original descriptors on St. Peter's Basilica
SuperPoint and DISK with SANDesc descriptors (top) vs. their original descriptors (bottom).
RIPE with SANDesc vs. original descriptors on Brandenburg Gate
RIPE with SANDesc descriptors (left) vs. its original descriptors (right).

TL;DR

SANDesc can be trained on top of any keypoint detector, leaving the detector untouched. It matches state-of-the-art descriptors on low-resolution data and surpasses all of them on high-resolution data, while still running on a single consumer GPU with 24 GB of VRAM.

Contributions

  • We propose SANDesc, a Streamlined Attention-based Network for Descriptor extraction that can be trained on top of any existing keypoint detector, improving matching without modifying the detector itself.
  • We design a revised U-Net-like architecture built from Residual U-Net Blocks with Attention (Convolutional Block Attention Modules plus residual paths), reaching strong local representation with just 2.4M parameters.
  • We train with a modified triplet loss combined with a curriculum-learning–inspired hard negative mining strategy, which improves training stability.
  • We show on HPatches, MegaDepth-1500, and the Image Matching Challenge 2021 that SANDesc is on par with state-of-the-art descriptors at low resolution, and we introduce a new urban 4K dataset with pre-calibrated intrinsics on which it substantially outperforms them — all while training and running on a single 24 GB consumer GPU.
RUBA
Residual U-Net Block with Attention (RUBA) structure.
Architecture
SANDesc architecture overview.

Quantitative Results

Graz4K @ 4K

MegaDepth-1500 @ 2048 kpts

IMC 2021 (avg over 9 scenes)

SuperPoint DISK RIPE ALIKED DeDoDe-G +SANDesc

AUC@5 (higher is better) across three matching benchmarks. DeDoDe-G runs out of memory on Graz4K at 4K.

Graz4K Sparse Reconstructions

Castle
Church
Clocktower
Main Square
Townhall
University

BibTeX

@inproceedings{durso2026sandesc,
        title={A Streamlined Attention-based Network for Descriptor Extraction},
        author={D'Urso, Mattia and Santellani, Emanuele and Sormann, Christian and Rossi, Mattia and Kuhn, Andreas and Fraundorfer, Friedrich},
        booktitle={2026 International Conference on 3D Vision (3DV)},
        year={2026},
        organization={IEEE Computer Society}
      }