Skip to main content
Published

Position Encoding in Detection-Based LiDAR-Camera Matching: A Diagnostic Study At Infrastructure Sites

Lihao Guo, Jiahao Tang, Tam Bang, Tianya Zhang, Austin Harris, Mina Sartipi, Siyang Cao

IEEE Sensors Letters, 2026 , Vol. 10 (8) , pp. 6007404

Abstract

We present a diagnostic study of learnable position encoding (PE) in the matching front-end of detection-based light detection and ranging (LiDAR)-camera calibration (detect, match, perspective-n-point). Across three infrastructure sites (A9 Highway, UTC3, and UTC4), four PE variants are evaluated on a ResNet backbone and a crop-level Vision Transformer. Learnable 2D-only position fusion reaches 97.9% rank-1 matching accuracy and matches or exceeds the full depth-aware variant at every site. Bidirectional UTC4-UTC3 zero-shot transfer collapses all architectures by 65%-93%, suggesting learnable position fusion may contribute to cross-site generalization failure together with appearance and domain shift. A two-stage Top-K retrieve-with-refine pipeline reduces matching latency from seconds to tens of milliseconds on a desktop GPU.

Citation

Lihao Guo, Jiahao Tang, Tam Bang, Tianya Zhang, Austin Harris, Mina Sartipi, Siyang Cao. "Position Encoding in Detection-Based LiDAR-Camera Matching: A Diagnostic Study At Infrastructure Sites." IEEE Sensors Letters 10(8): 6007404, 2026. DOI: 10.1109/LSENS.2026.3708462

Publication Snapshot

Authors
Lihao Guo, Jiahao Tang, Tam Bang, Tianya Zhang, Austin Harris, Mina Sartipi, Siyang Cao
Venue
IEEE Sensors Letters
Published
2026
Pages
6007404

Quick Links

Details

Year
2026
Published In
IEEE Sensors Letters
Volume
10
Issue
8
Pages
6007404
DOI
10.1109/LSENS.2026.3708462