arXiv ScienceSearch

arXiv · 2512.02972

BEVDilation: LiDAR-Centric Multi-Modal Fusion for 3D Object Detection

Abstract

Integrating LiDAR and camera information in the bird's eye view (BEV) representation has demonstrated its effectiveness in 3D object detection. However, because of the fundamental disparity in geometric accuracy between these sensors, indiscriminate fusion in previous methods often leads to degraded performance. In this paper, we propose BEVDilation, a novel LiDAR-centric framework that prioritizes LiDAR information in the fusion. By formulating image BEV features as implicit guidance rather than naive concatenation, our strategy effectively alleviates the spatial misalignment caused by image depth estimation errors. Furthermore, the image guidance can effectively help the LiDAR-centric paradigm to address the sparsity and semantic limitations of point clouds. Specifically, we propose a Sparse Voxel Dilation Block that mitigates the inherent point sparsity by densifying foreground voxels through image priors. Moreover, we introduce a Semantic-Guided BEV Dilation Block to enhance the LiDAR feature diffusion processing with image semantic guidance and long-range context capture. On the challenging nuScenes benchmark, BEVDilation achieves better performance than state-of-the-art methods while maintaining competitive computational efficiency. Importantly, our LiDAR-centric strategy demonstrates greater robustness to depth noise compared to naive fusion. The source code is available at https://github.com/gwenzhang/BEVDilation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Guowen Zhang, Chenhang He, Liyi Chen, Lei Zhang. 2025-12-02. BEVDilation: LiDAR-Centric Multi-Modal Fusion for 3D Object Detection. https://arxiv.org/abs/2512.02972

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

RealLiFe: Real-Time Light Field Reconstruction via Hierarchical Sparse Gradient Descent

With the rise of Extended Reality (XR) technology, there is a growing need for real-time light field reconstruction from sparse view inputs. Existing methods can be classified into offline techniques, which can generate high-quality novel views but at the cost of long inference/training time, and online methods, which either lack generalizability or produce unsatisfactory results. However, we have observed that the intrinsic sparse manifold of Multi-plane Images (MPI) enables a significant acceleration of light field reconstruction while maintaining rendering quality. Based on this insight, we introduce RealLiFe, a novel light field optimization method, which leverages the proposed Hierarchical Sparse Gradient Descent (HSGD) to produce high-quality light fields from sparse input images in real time. Technically, the coarse MPI of a scene is first generated using a 3D CNN, and it is further optimized leveraging only the scene content aligned sparse MPI gradients in a few iterations. Extensive experiments demonstrate that our method achieves comparable visual quality while being 100x faster on average than state-of-the-art offline methods and delivers better performance (about 2 dB higher in PSNR) compared to other online approaches.

cs.CV

SegCol Challenge: Semantic Segmentation for Tools and Fold Edges in Colonoscopy data

Improving the reliability and completeness of colonoscopic inspection is critical for reducing missed lesions and improving colorectal cancer prevention. Reliable scene understanding is essential for navigation, reconstruction, and assessment of inspection completeness. Anatomical structures such as mucosal folds provide stable geometric cues for endoscope localization, while surgical instruments introduce dynamic occlusions that complicate visual interpretation. However, existing gastrointestinal endoscopy datasets largely focus on disease detection or artifact segmentation, leaving a gap in precise annotations of structural landmarks and instruments. We introduce SegCol, a dataset and benchmark for semantic segmentation of colon fold edges and surgical instruments derived from the EndoMapper dataset. SegCol provides manually annotated pixel-level masks for three instrument classes and thin fold-edge structures across temporally consistent image sequences. It forms the basis of the SegCol Challenge, organized as part of the EndoVis Challenge at MICCAI 2024, evaluating both supervised segmentation and annotation-efficient active learning. We further study segmentation metrics, including Dice, ODS/OIS, AP, and CLDice, under structural perturbations and different object geometries, and analyze participating methods, architectural choices, and active learning strategies. Our findings show that metric behavior strongly depends on target structure, highlighting the need for carefully selected evaluation protocols in endoscopic segmentation. Details are available at https://www.synapse.org/Synapse:syn54124209/wiki/626563, and code at https://github.com/surgical-vision/segcol_challenge.

cs.CV

evMLP: An Efficient Event-Driven MLP Architecture for Vision

While CNNs and ViTs dominate vision architectures, all-MLP models offer a structurally simpler alternative whose patch-independent processing is naturally suited to exploiting temporal redundancy in video. We present evMLP, an all-MLP architecture that processes image patches independently, enabling an event-driven local update mechanism for video processing: by defining inter-frame changes as "events" and processing only the patches where events occur, evMLP avoids redundant computation on unchanged regions. Because each patch is processed independently, skipping an unchanged patch leaves all other outputs unaffected; at an event threshold of zero, the mechanism produces outputs identical to the dense baseline rather than an approximation. On ImageNet, evMLP achieves 73.5% top-1 accuracy at 1.03 GMACs (rising to 77.0% with knowledge distillation and an extended training schedule). On multiple video datasets, the event-driven mechanism reduces computational cost by 8.4%-26.8% while maintaining output consistency with the dense baseline. Wall-clock measurements confirm that these savings translate into actual speedup under compute-bound conditions, and that stream-level parallelism is the effective deployment strategy for multi-core systems. The code and pre-trained models are available at https://github.com/i-evi/evMLP.

cs.CV