Abstract:In order to solve the problem that it is difficult to effectively fuse point cloud data with image and video data in the current mainstream bird’s eye view fusion (BEVFusion) method, a DC fusion module based on deformable attention mechanism and cross attention mechanism is proposed. Through the fusion of deformable attention mechanism and cross-attention mechanism, the target point cloud and image features are enhanced, and the errors caused by fusion are reduced. The experimental results on nuScenes show that the average detection accuracy of DC-BEVFusion is 65.6%, which is 6.9% higher than that of BEVFusion, and the detection results are more accurate and robust.