(Full) Attention is More Than You Need: Sparse Cross-View Attention in VGGT 馃敆
Semester Project 路 EPFL CVLab (2026)
Replaced VGGT's dense global cross-view attention with a learned sparse index over patch correspondences: 32x faster than dense FlashAttention at 96 frames, and gains +5 pose AUC@30 on ScanNet. A contrastive InfoNCE loss teaches the backbone to build its own sparse index at inference, with no geometry.







