Position Encoding in Detection-Based LiDAR-Camera Matching: A Diagnostic Study At Infrastructure Sites
Lihao Guo, Jiahao Tang, Tam Bang, Tianya Zhang, Austin Harris, Mina Sartipi, Siyang Cao
IEEE Sensors Letters, 2026 , Vol. 10 (8) , pp. 6007404
Abstract
We present a diagnostic study of learnable position encoding (PE) in the matching front-end of detection-based light detection and ranging (LiDAR)-camera calibration (detect, match, perspective-n-point). Across three infrastructure sites (A9 Highway, UTC3, and UTC4), four PE variants are evaluated on a ResNet backbone and a crop-level Vision Transformer. Learnable 2D-only position fusion reaches 97.9% rank-1 matching accuracy and matches or exceeds the full depth-aware variant at every site. Bidirectional UTC4-UTC3 zero-shot transfer collapses all architectures by 65%-93%, suggesting learnable position fusion may contribute to cross-site generalization failure together with appearance and domain shift. A two-stage Top-K retrieve-with-refine pipeline reduces matching latency from seconds to tens of milliseconds on a desktop GPU.
Citation
Lihao Guo, Jiahao Tang, Tam Bang, Tianya Zhang, Austin Harris, Mina Sartipi, Siyang Cao. "Position Encoding in Detection-Based LiDAR-Camera Matching: A Diagnostic Study At Infrastructure Sites." IEEE Sensors Letters 10(8): 6007404, 2026. DOI: 10.1109/LSENS.2026.3708462
Publication Snapshot
- Authors
- Lihao Guo, Jiahao Tang, Tam Bang, Tianya Zhang, Austin Harris, Mina Sartipi, Siyang Cao
- Venue
- IEEE Sensors Letters
- Published
- 2026
- Pages
- 6007404
Quick Links
Details
- Year
- 2026
- Published In
- IEEE Sensors Letters
- Volume
- 10
- Issue
- 8
- Pages
- 6007404
- DOI
- 10.1109/LSENS.2026.3708462