Dong, Guirong, Wu, Yi, Dai, Leyu, Chang, Han, Tang, Zhaoxun and Liu, Dianzi (2026) LDFpose: A lightweight deep fusion network for 6D occlusion object pose estimation. Image and Vision Computing, 174. ISSN 0262-8856
Preview |
PDF (Revised_version_for_acceptance_version)
- Accepted Version
Available under License Creative Commons Attribution Non-commercial No Derivatives. Download (1MB) | Preview |
Abstract
Effectively implementing color and depth modalities into the model for accurate 6D object pose estimation in the presence of occlusions and complex backgrounds remains a challenge in the field of robotics vision. To tackle these problems, a novel method LDFpose, a lightweight deep fusion network based on RGB-D input is proposed for real-time and precise object pose estimation in this paper. Unlike existing methods, this developed method effectively captures global semantic information from complex textured backgrounds for the color feature extraction using an Embedding-Free Hierarchical Transformer (EHT) encoder, while significantly reducing computational overhead. In the point cloud feature extraction, an Adaptive Hierarchical Keypoints (AHK) sampling strategy is applied, which employs sparse sampling of distant key points to expand the receptive field and improve the efficiency of point cloud interactions. Furthermore, the model adaptability to heterogeneous inputs and its global representation capability are enhanced by integrating RGB and geometric features in use of the Adaptive Modal Fusion (AMF) block, facilitating more effective complementary data fusion. Experimental results on the datasets including Linemod, YCB-Video, Occlusion Linemod, and a custom dataset of Occlusion of Apes (OccoA) demonstrate that LDFpose outperforms existing methods in pose estimation accuracy and inference speed, particularly in scenarios with different lighting conditions and high degrees of occlusion arising from practical industrial applications.
| Item Type: | Article |
|---|---|
| Additional Information: | Data availability: Data will be made available on request. |
| Uncontrolled Keywords: | 6d pose estimation,adaptive hierarchical keypoints sampling,cross-modality attention fusion,embedding-free hierarchical attention,signal processing,computer vision and pattern recognition ,/dk/atira/pure/subjectarea/asjc/1700/1711 |
| Faculty \ School: | Faculty of Science > School of Engineering, Mathematics and Physics |
| UEA Research Groups: | Faculty of Science > Research Groups > Sustainable Energy |
| Related URLs: | |
| Depositing User: | LivePure Connector |
| Date Deposited: | 13 Aug 2026 14:17 |
| Last Modified: | 13 Aug 2026 14:17 |
| URI: | https://ueaeprints.uea.ac.uk/id/eprint/104119 |
| DOI: | 10.1016/j.imavis.2026.106147 |
Downloads
Downloads per month over past year
Actions (login required)
![]() |
View Item |
Tools
Tools