1.青岛科技大学数据科学学院,山东青岛 266061
2.中国海洋大学信息科学与工程学部,山东青岛 266100
3.荷兰埃因霍芬理工大学工业工程学院,荷兰 5612 AE
樊怡麟 男,2002年5月出生于山东省淄博市。2026年毕业于青岛科技大学数据科学学院。现为软控股份有限公司算法工程师。主要研究方向为大模型应用、开放词汇多目标检测。E-mail: fylqust@163.com
杨方正 男,2003年7月出生于山东省滨州市。现为青岛科技大学硕士研究生。主要研究方向为多目标检测与跟踪。E-mail: fzyang@mails.qust.edu.cn
臧银珠 女,2003年3月出生于山东省日照市。现为青岛科技大学硕士研究生。主要研究方向为计算机视觉、三维建模。E-mail: zangyinzhu@mails.qust.edu.cn
李辉 男,1984年3月出生于河南省平顶山市。现为青岛科技大学数据科学学院副教授、硕士生导师。主要研究方向为计算机视觉,多目标检测与跟踪。E-mail: lihui@qust.edu.cn
刘治宇 男,1988年3月出生于黑龙江省七台河市。现为中国海洋大学博士后。主要研究方向为多模态特征编码、AI高性能推理计算等方向。E-mail: lzy1280@ouc.edu.cn
孔祥振 男,1990年12月出生于山东省聊城市。现为荷兰埃因霍芬理工大学工业工程学院博士后。主要研究方向为图像处理和目标检测。E-mail: x.kong@tue.nl
收稿:2026-07-01,
录用:2026-07-13,
网络首发:2026-08-06,
移动端阅览
樊怡麟, 杨方正, 臧银珠, 等. 基于空间重构和文本引导的室内开放域3D目标检测方法[J/OL]. 电子学报, 2026,1-12.
FAN Yilin, YANG Fangzheng, ZANG Yinzhu, et al. Indoor Open Domain 3D Object Detection Method Based on Spatial Reconstruction and Text Guidance[J/OL]. ACTA ELECTRONICA SINICA, 2026, 1-12.
樊怡麟, 杨方正, 臧银珠, 等. 基于空间重构和文本引导的室内开放域3D目标检测方法[J/OL]. 电子学报, 2026,1-12. DOI: 10.12263/DZXB.20260414.
FAN Yilin, YANG Fangzheng, ZANG Yinzhu, et al. Indoor Open Domain 3D Object Detection Method Based on Spatial Reconstruction and Text Guidance[J/OL]. ACTA ELECTRONICA SINICA, 2026, 1-12. DOI: 10.12263/DZXB.20260414.
3D目标检测是提升智能系统在复杂室内环境下自主决策与精细化感知能力的关键。传统视觉方法受室内物体长尾分布影响显著,且缺乏文本语义的引导,导致模型难以对室内场景中不断涌现的未见类别进行有效泛化。针对室内复杂场景下目标类别繁杂且分布不均导致的多尺度语义混叠问题,以及跨模态异构空间中由于背景噪声引发的表征分布失配挑战,本文提出一种基于空间重构和文本引导的室内开放域3D目标检测方法。在视觉表征前端,本文设计层级边缘聚焦网络,利用分流注意力机制过滤通道冗余并增强跨尺度语义一致性,构建层级融合模块生成自适应层间注意力,从而有效解决由尺度变化引发的检测歧义。在多模态融合阶段,本文构建基于空间重构的跨模态特征融合模块,依据统计特性将异构特征空间解耦为信息流与冗余流,通过提纯高价值视觉特征缓解室内杂乱环境引起的信噪比差异。在解码阶段,本文建立语义增强的文本查询交互机制,将对比语言-图像预训练(Contrastive Language-Image Pre-training,CLIP)文本嵌入作为先验注入解码器,赋予查询向量目标导向的感知能力,降低细粒度类别的识别误差。在SUN RGB-D数据集上的实验结果表明,该方法在新颖类别上的检测精度达11.33%,全类别平均检测精度达19.30%,较基线方法分别提升了1.32%和1.15%,证明了该方法在室内开放域场景下具有优越的3D目标检测性能。
3D object detection is crucial for enhancing the autonomous decision-making and refined perception capabilities of intelligent systems in complex indoor environments. Traditional visual methods are significantly affected by the long-tail distribution of indoor objects and lack textual semantic guidance
making it difficult for models to effectively generalize to the constantly emerging unseen categories within indoor scenes. To address the challenges of multi-scale semantic aliasing caused by the complex and uneven distribution of object categories in complex indoor scenes
and the representational distribution mismatch caused by background noise in cross-modal heterogeneous spaces
this paper proposes an indoor open-domain 3D object detection method based on spatial reconstruction and textual guidance. At the visual representation front end
this paper designs a hierarchical edge-focusing network
utilizing a shunt attention mechanism to filter channel redundancy and enhance cross-scale semantic consistency. This paper constructs a hierarchical fusion module to generate adaptive inter-layer attention
effectively resolving detection ambiguities caused by scale variations. In the multimodal fusion stage
this paper constructs a cross-modal feature fusion module based on spatial reconstruction. Based on statistical characteristics
the heterogeneous feature space is decoupled into information flow and redundant flow
and high-value visual features are purified to alleviate the signal-to-noise ratio differences caused by the cluttered indoor environment. In the decoding stage
this paper establishes a semantically enhanced text query interaction mechanism
which injects contrastive language-image pre-training (CLIP) text embeddings into the decoder as priors
endowing the query vectors with target-oriented perception capabilities and reducing recognition errors for fine-grained categories. Experimental results on the SUN RGB-D dataset show that the proposed method achieves a detection accuracy of 11.33% for novel categories and an average detection accuracy of 19.30% for all categories
representing improvements of 1.32% and 1.15% respectively compared to the baseline method. This demonstrates the superior 3D object detection performance of the proposed method in indoor open-domain scenes.
Lazarow J , Griffiths D , Kohavi G , et al . Cubify anything: Scaling indoor 3D object detection [C ] // 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2025 : 22225 - 22233 . DOI: 10.1109/cvpr52734.2025.02070 http://dx.doi.org/10.1109/cvpr52734.2025.02070
Wang Jiangyi , Zhao Na . Uncertainty meets diversity: A comprehensive active learning framework for indoor 3D object detection [C ] // 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2025 : 20329 - 20339 . DOI: 10.1109/cvpr52734.2025.01893 http://dx.doi.org/10.1109/cvpr52734.2025.01893
Zhu Chaoyang , Chen Long . A survey on open-vocabulary detection and segmentation: Past, present, and future [J ] . IEEE Transactions on Pattern Analysis and Machine Intelligence , 2024 , 46 ( 12 ): 8954 - 8975 . DOI: 10.1109/tpami.2024.3413013 http://dx.doi.org/10.1109/tpami.2024.3413013
Li Kaiyu , Cao Xiangyong , Deng Yupeng , et al . DynamicEarth: How far are we from open-vocabulary change detection [J ] . Proceedings of the AAAI Conference on Artificial Intelligence , 2026 , 40 ( 8 ): 6279 - 6287 . DOI: 10.1609/aaai.v40i8.37554 http://dx.doi.org/10.1609/aaai.v40i8.37554
Hu Yupeng , Ding Changxing , Sun Chang , et al . Bilateral collaboration with large vision-language models for open vocabulary human-object interaction detection [C ] // 2025 IEEE/CVF International Conference on Computer Vision . Piscataway : IEEE , 2025 : 20126 - 20136 . DOI: 10.1109/iccv51701.2025.01872 http://dx.doi.org/10.1109/iccv51701.2025.01872
Radford A , Kim J W , Hallacy C , et al . Learning transferable visual models from natural language supervision [C ] // International Conference on Machine Learning . 2021 : 8748 - 8763 . DOI: 10.48550/arXiv.2103.00020 http://dx.doi.org/10.48550/arXiv.2103.00020
周治国 , 马文浩 . 一种多层多模态融合3D目标检测方法 [J ] . 电子学报 , 2024 , 52 ( 3 ): 696 - 708 . DOI: 10.12263/DZXB.20220593 http://dx.doi.org/10.12263/DZXB.20220593
Zhou Zhiguo , Ma Wenhao . 3D object detection based on multilayer multimodal fusion [J ] . Acta Electronica Sinica , 2024 , 52 ( 3 ): 696 - 708 . (in Chinese) . DOI: 10.12263/DZXB.20220593 http://dx.doi.org/10.12263/DZXB.20220593
Xu Danfei , Anguelov D , Jain A . PointFusion: Deep sensor fusion for 3D bounding box estimation [C ] // 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2018 : 244 - 253 . DOI: 10.1109/cvpr.2018.00033 http://dx.doi.org/10.1109/cvpr.2018.00033
Qi C R , Chen Xinlei , Litany O , et al . ImVoteNet: Boosting 3D object detection in point clouds with image votes [C ] // 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2020 : 4403 - 4412 . DOI: 10.1109/cvpr42600.2020.00446 http://dx.doi.org/10.1109/cvpr42600.2020.00446
葛同澳 , 李辉 , 郭颖 等 . 基于双融合框架的多模态3D目标检测算法 [J ] . 电子学报 , 2023 , 51 ( 11 ): 3100 - 3110 . DOI: 10.12263/DZXB.20230414 http://dx.doi.org/10.12263/DZXB.20230414 .
GE Tongao , LI Hui , GUO Ying , et al . A Multimodal 3D Object Detection Method Based on Double-Fusion Framework [J ] . ACTA ELECTRONICA SINICA , 2023 , 51 ( 11 ): 3100 - 3110 . DOI: 10.12263/DZXB.20230414. http://dx.doi.org/10.12263/DZXB.20230414. (in Chinese)
Vora S , Lang A H , Helou B , et al . PointPainting: Sequential fusion for 3D object detection [C ] // 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2020 : 4603 - 4611 . DOI: 10.1109/cvpr42600.2020.00466 http://dx.doi.org/10.1109/cvpr42600.2020.00466
Zhang Zehan , Zhang Ming , Liang Zhidong , et al . MAFF-net: Filter false positive for 3D vehicle detection with multi-modal adaptive feature fusion [C ] // 2022 IEEE 25th International Conference on Intelligent Transportation Systems . Piscataway : IEEE , 2022 : 369 - 376 . DOI: 10.1109/itsc55140.2022.9922104 http://dx.doi.org/10.1109/itsc55140.2022.9922104
Tan Xun , Chen Xingyu , Zhang Guowei , et al . MBDF-net: Multi-branch deep fusion network for 3D object detection [C ] // Proceedings of the 1st International Workshop on Multimedia Computing for Urban Data . New York : ACM , 2021 : 9 - 17 . DOI: 10.1145/3475721.3484311 http://dx.doi.org/10.1145/3475721.3484311
Wang Zhixin , Jia Kui . Frustum ConvNet: Sliding Frustums to aggregate local point-wise features for amodal 3D object detection [C ] // 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems . Piscataway : IEEE , 2019 : 1742 - 1749 . DOI: 10.1109/iros40897.2019.8968513 http://dx.doi.org/10.1109/iros40897.2019.8968513
臧传方 , 党建武 , 雍玖 . 自适应融合多模态特征的6D物体位姿估计方法 [J ] . 激光与光电子学进展 , 2025 , 62 ( 4 ): 0415002 .
Zang Chuanfang , Dang Jianwu , Yong Jiu . Adaptive multimodal-feature fusion for 6D object position estimation [J ] . Laser & Optoelectronics Progress , 2025 , 62 ( 4 ): 0415002 . (in Chinese)
Gu Xiuye , Lin Tsung-Yi , Kuo Weicheng , et al . Open-vocabulary object detection via vision and language knowledge distillation [PP/OL ] . V3. arXiv ( 2022-05-12 )[ 2026-07-01 ] . https://doi.org/10.48550/arXiv.2104.13921 https://doi.org/10.48550/arXiv.2104.13921 .
Minderer M , Gritsenko A , Stone A , et al . Simple open-vocabulary object detection [C ] // Computer Vision - ECCV 2022 . Cham : Springer , 2022 : 728 - 755 . DOI: 10.1007/978-3-031-20080-9_42 http://dx.doi.org/10.1007/978-3-031-20080-9_42
樊琳 , 龚勋 , 郑岑洋 . 基于文本引导下的多模态医学图像分析算法 [J ] . 电子学报 , 2024 , 52 ( 7 ): 2341 - 2355 .
Fan Lin , Gong Xun , Zheng Cenyang . A multi-modal medical image analysis algorithm based on text guidance [J ] . Acta Electronica Sinica , 2024 , 52 ( 7 ): 2341 - 2355 . (in Chinese)
Cheng Tianheng , Song Lin , Ge Yixiao , et al . YOLO-world: Real-time open-vocabulary object detection [C ] // 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2024 : 16901 - 16911 . DOI: 10.1109/cvpr52733.2024.01599 http://dx.doi.org/10.1109/cvpr52733.2024.01599
Liu Shilong , Zeng Zhaoyang , Ren Tianhe , et al . Grounding DINO: Marrying DINO with Grounded pre-training for Open-set object detection [C ] // Computer Vision - ECCV 2024 . Cham : Springer , 2025 : 38 - 55 . DOI: 10.1007/978-3-031-72970-6_3 http://dx.doi.org/10.1007/978-3-031-72970-6_3
Wang Zhenyu , Li Yali , Liu Taichi , et al . OV-Uni3DETR: Towards unified open-vocabulary 3D object detection viaCycle-modality propagation [C ] // Computer Vision - ECCV 2024 . Cham : Springer , 2025 : 73 - 89 . DOI: 10.1007/978-3-031-72970-6_5 http://dx.doi.org/10.1007/978-3-031-72970-6_5
Zhang Renrui , Guo Ziyu , Zhang Wei , et al . PointCLIP: Point cloud understanding by CLIP [C ] // 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2022 : 8542 - 8552 . DOI: 10.1109/cvpr52688.2022.00836 http://dx.doi.org/10.1109/cvpr52688.2022.00836
Zhu Xiangyang , Zhang Renrui , He Bowei , et al . PointCLIP V2: Prompting CLIP and GPT for powerful 3D open-world learning [C ] // 2023 IEEE/CVF International Conference on Computer Vision . Piscataway : IEEE , 2023 : 2639 - 2650 . DOI: 10.1109/iccv51070.2023.00249 http://dx.doi.org/10.1109/iccv51070.2023.00249
Zeng Yihan , Jiang Chenhan , Mao Jiageng , et al . CLIP2: Contrastive language-image-point pretraining from real-world point cloud data [C ] // 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2023 : 15244 - 15253 . DOI: 10.1109/cvpr52729.2023.01463 http://dx.doi.org/10.1109/cvpr52729.2023.01463
Cao Yang , Zeng Yihan , Xu Hang , et al . CoDA: Collaborative novel box discovery and cross-modal alignment for open-vocabulary 3D object detection [C ] // Advances in Neural Information Processing Systems 36 . Neural Information Processing Systems Foundation, Inc. (NeurIPS) , 2023 : 71862 - 71873 . DOI: 10.52202/075280-3145 http://dx.doi.org/10.52202/075280-3145
Jiao Pengkun , Zhao Na , Chen Jingjing , et al . Unlocking textual and visual wisdom: Open-vocabulary 3D object detection enhanced by comprehensive guidance from text and image [C ] // Computer Vision - ECCV 2024 . Cham : Springer , 2025 : 376 - 392 . DOI: 10.1007/978-3-031-73195-2_22 http://dx.doi.org/10.1007/978-3-031-73195-2_22
Cao Yang , Zeng Yihan , Xu Hang , et al . Collaborative novel object discovery and box-guided cross-modal alignment for open-vocabulary 3D object detection [J ] . IEEE Transactions on Pattern Analysis and Machine Intelligence , 2025 , 47 ( 11 ): 10475 - 10489 . DOI: 10.1109/tpami.2025.3593580 http://dx.doi.org/10.1109/tpami.2025.3593580
Peng Xingyu , Liu Si , Gao Chen , et al . GLRD: Global-local collaborative reason and debate with PSL for 3D open-vocabulary detection [PP/OL ] . V1. arXiv ( 2025-03-26 )[ 2026-07-01 ] . https://doi.org/10.48550/arXiv.2503.20682 https://doi.org/10.48550/arXiv.2503.20682 .
0
浏览量
0
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621