LDFpose: A lightweight deep fusion network for 6D occlusion object pose estimation

Dong, Guirong, Wu, Yi, Dai, Leyu, Chang, Han, Tang, Zhaoxun and Liu, Dianzi (2026) LDFpose: A lightweight deep fusion network for 6D occlusion object pose estimation. Image and Vision Computing, 174. ISSN 0262-8856

[thumbnail of Revised_version_for_acceptance_version]
Preview
PDF (Revised_version_for_acceptance_version) - Accepted Version
Available under License Creative Commons Attribution Non-commercial No Derivatives.

Download (1MB) | Preview

Abstract

Effectively implementing color and depth modalities into the model for accurate 6D object pose estimation in the presence of occlusions and complex backgrounds remains a challenge in the field of robotics vision. To tackle these problems, a novel method LDFpose, a lightweight deep fusion network based on RGB-D input is proposed for real-time and precise object pose estimation in this paper. Unlike existing methods, this developed method effectively captures global semantic information from complex textured backgrounds for the color feature extraction using an Embedding-Free Hierarchical Transformer (EHT) encoder, while significantly reducing computational overhead. In the point cloud feature extraction, an Adaptive Hierarchical Keypoints (AHK) sampling strategy is applied, which employs sparse sampling of distant key points to expand the receptive field and improve the efficiency of point cloud interactions. Furthermore, the model adaptability to heterogeneous inputs and its global representation capability are enhanced by integrating RGB and geometric features in use of the Adaptive Modal Fusion (AMF) block, facilitating more effective complementary data fusion. Experimental results on the datasets including Linemod, YCB-Video, Occlusion Linemod, and a custom dataset of Occlusion of Apes (OccoA) demonstrate that LDFpose outperforms existing methods in pose estimation accuracy and inference speed, particularly in scenarios with different lighting conditions and high degrees of occlusion arising from practical industrial applications.

Item Type: Article
Additional Information: Data availability: Data will be made available on request.
Uncontrolled Keywords: 6d pose estimation,adaptive hierarchical keypoints sampling,cross-modality attention fusion,embedding-free hierarchical attention,signal processing,computer vision and pattern recognition ,/dk/atira/pure/subjectarea/asjc/1700/1711
Faculty \ School: Faculty of Science > School of Engineering, Mathematics and Physics
UEA Research Groups: Faculty of Science > Research Groups > Sustainable Energy
Related URLs:
Depositing User: LivePure Connector
Date Deposited: 13 Aug 2026 14:17
Last Modified: 13 Aug 2026 14:17
URI: https://ueaeprints.uea.ac.uk/id/eprint/104119
DOI: 10.1016/j.imavis.2026.106147

Downloads

Downloads per month over past year

Actions (login required)

View Item View Item