arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

自动驾驶

自动驾驶感知、规划、BEV、占用预测、激光雷达和仿真评测。

2026-03-10 至 2026-03-10 共收录 17 信号源:cs.RO, cs.CV, eess.IV, cs.AI

1. 感知 17 篇

2603.08180 2026-03-10 cs.CV cs.LG 83%

ALOOD: Exploiting Language Representations for LiDAR-based Out-of-Distribution Object Detection

ALOOD:利用语言表示进行基于LiDAR的分布外物体检测

Michael Kösel, Marcel Schreiber, Michael Ulrich, Claudius Gläser, Klaus Dietmayer

专题命中 感知 :LiDAR(title,abstract);autonomous driving(abstract);分类 cs.CV

AI总结 ALOOD通过整合语言表示,提出了一种基于LiDAR的分布外物体检测新方法,利用视觉-语言模型实现零样本分类。

Comments Accepted for publication at the 2025 IEEE Intelligent Transportation Systems Conference (ITSC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07593 2026-03-10 cs.CV 83%

Fast Attention-Based Simplification of LiDAR Point Clouds for Object Detection and Classification

基于快速注意力机制的激光雷达点云简化用于目标检测与分类

Z. Rozsa, Á. Madaras, Q. Wei, X. Lu, M. Golarits, H. Yuan, T. Sziranyi, R. Hamzaoui

机构 * Institute for Computer Science and Control (SZTAKI)(计算机科学与控制研究所) School of Information Engineering(信息工程学院) Faculty of Technology, Arts, and Culture(技术、艺术与文化学院) School of Control Science and Engineering(控制科学与工程学院)

专题命中 感知 :LiDAR(title,abstract);autonomous driving(abstract);分类 cs.CV

AI总结 本文提出一种基于注意力机制的高效激光雷达点云简化方法,通过特征嵌入与注意力采样模块提升目标检测与分类的效率和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08254 2026-03-10 cs.CV 79%

DynamicVGGT: Learning Dynamic Point Maps for 4D Scene Reconstruction in Autonomous Driving

DynamicVGGT: 为自动驾驶中的4D场景重建学习动态点图

Zhuolin He, Jing Li, Guanghao Li, Xiaolei Chen, Jiacheng Tang, Siyang Zhang, Zhounan Jin, Feipeng Cai, Bin Li, Jian Pu, Jia Cai, Xiangyang Xue

机构 * Fudan University(复旦大学) Huawei(华为) Yinwang Intelligent Technology(依文智能技术) Shanghai Innovation Institute(上海创新研究院) CUHK(香港中文大学)

专题命中 感知 :autonomous driving(title,abstract);分类 cs.CV

AI总结 DynamicVGGT通过联合预测当前和未来点图,结合Motion-aware Temporal Attention模块和动态3D高斯点散射头,实现了自动驾驶中4D动态场景的高精度重建。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07493 2026-03-10 cs.CV 77%

RayD3D: Distilling Depth Knowledge Along the Ray for Robust Multi-View 3D Object Detection

RayD3D: 沿射线蒸馏深度知识以实现鲁棒的多视角3D目标检测

Rui Ding, Zhaonian Kuang, Zongwei Zhou, Meng Yang, Xinhu Zheng, Gang Hua

专题命中 感知 :autonomous driving(abstract);BEV(abstract);LiDAR(abstract);分类 cs.CV

AI总结 RayD3D通过沿射线蒸馏深度知识,提升多视角3D目标检测的鲁棒性,无需增加计算成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08611 2026-03-10 cs.CV cs.RO 73%

FOMO-3D: Using Vision Foundation Models for Long-Tailed 3D Object Detection

FOMO-3D:利用视觉基础模型进行长尾3D目标检测

Anqi Joyce Yang, James Tu, Nikita Dvornik, Enxu Li, Raquel Urtasun

机构 * Waabi University of Toronto(多伦多大学)

专题命中 感知 :self-driving(abstract);LiDAR(abstract);分类 cs.RO、cs.CV

AI总结 FOMO-3D通过利用视觉基础模型的丰富先验知识和多模态融合设计,提升长尾3D目标检测的性能。

Comments Published at 9th Annual Conference on Robot Learning (CoRL 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07486 2026-03-10 cs.CV 70%

Multi-Modal Decouple and Recouple Network for Robust 3D Object Detection

多模态解耦与耦合网络用于抗干扰的3D目标检测

Rui Ding, Zhaonian Kuang, Yuzhe Ji, Meng Yang, Xinhu Zheng, Gang Hua

机构 * State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(人机混合增强智能国家重点实验室,人工智能与机器人研究院,西安交通大学) Intelligent Transportation Thrust of the Systems Hub, The Hong Kong University of Science and Technology (Guangzhou)(系统枢纽智能交通方向,香港科技大学(广州)) Multimodal Experiences Research Lab, Dolby Laboratories(多模态体验研究实验室,Dolby实验室)

专题命中 感知 :BEV(abstract);LiDAR(abstract);分类 cs.CV

AI总结 本文提出多模态解耦与耦合网络,通过分离和重新耦合不同模态特征以提高在数据损坏下的3D目标检测鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19195 2026-03-10 cs.CV cs.AI 62%

Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks

重新思考驾驶世界模型作为感知任务的合成数据生成器

Kai Zeng, Zhanqian Wu, Kaixin Xiong, Xiaobao Wei, Xiangyu Guo, Zhenxin Zhu, Kalok Ho, Lijun Zhou, Bohan Zeng, Ming Lu, Haiyang Sun, Bing Wang, Guang Chen, Hangjun Ye, Wentao Zhang

机构 * Peking University(北京大学) Xiaomi EV(小米电动车) Huazhong University of Science and Technology(华中科技大学) Beijing Key Laboratory of Data Intelligence and Security (Peking University)(北京数据智能与安全重点实验室(北京大学)) Zhongguancun Academy(中关村学院)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV、cs.AI

AI总结 Dream4Drive通过生成高质量的合成数据提升自动驾驶感知任务性能

Journal ref ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13912 2026-03-10 cs.RO cs.AI 62%

Energy-Efficient SLAM via Joint Design of Sensing, Communication, and Exploration Speed

通过联合设计感知、通信和探索速度实现节能SLAM

Zidong Han, Ruibo Jin, Xiaoyang Li, Bingpeng Zhou, Qinyu Zhang, Yi Gong

机构 * Southern University of Science and Technology(南方科技大学) Harbin Institute of Technoglogy (Shenzhen)(哈尔滨工业大学(深圳)) Shenzhen Research Institute of Big Data(深圳大数据研究院) The School of Electronics and Communication Engineering(电子与通信工程学院) Sun Yat-sen University(中山大学)

专题命中 感知 :LiDAR(abstract);分类 cs.RO、cs.AI

AI总结 本文提出通过联合优化感知、通信和探索速度,实现节能的持续SLAM系统,利用深度学习方法提升地图重建效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08379 2026-03-10 cs.RO 57%

Perception-Aware Communication-Free Multi-UAV Coordination in the Wild

野外多无人机协同的感知-aware 通信-free 方法

Manuel Boldrer, Michal Kamler, Afzal Ahmad, Martin Saska

机构 * Department of Cybernetics, Czech Technical University in Prague(控制系,捷克技术大学)

专题命中 感知 :LiDAR(abstract);分类 cs.RO

AI总结 本研究提出了一种无需通信的多无人机协同方法,利用3D激光雷达进行SLAM和障碍物检测,实现复杂环境中的安全导航。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08240 2026-03-10 cs.CV 57%

SiMO: Single-Modality-Operable Multimodal Collaborative Perception

SiMO: 单模态可操作的多模态协作感知

Jiageng Wen, Shengjie Zhao, Bing Li, Jiafeng Huang, Kenan Ye, Hao Deng

机构 * Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University(同济大学智能自主系统上海研究院) School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院) School of Mechatronic Engineering and Automation, Shanghai University(上海大学机械电子工程与自动化学院)

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 SiMO通过单模态可操作的多模态协作感知方法,解决多模态特征融合中的语义不匹配问题,提升协作感知性能。

Comments Accepted to ICLR 2026. This arXiv version includes an additional appendix (Appendix 15) containing further philosophical discussion not included in the official ICLR peer-reviewed version

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00661 2026-03-10 cs.CV 57%

Elytra: A Flexible Framework for Securing Large Vision Systems

ELYTRA:一种用于安全大型视觉系统的灵活框架

Richard E. Neddo, Emmanuel Atindama, Zander W. Blasingame, Chen Liu

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 ELYTRA框架通过低秩适应技术,为大型视觉系统提供动态安全修补,有效提升对抗性示例下的分类准确率

Comments Updated pre-print. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07985 2026-03-10 cs.CV 57%

On the Feasibility and Opportunity of Autoregressive 3D Object Detection

关于自回归3D目标检测的可行性与机会

Zanming Huang, Jinsu Yoo, Sooyoung Jeon, Zhenzhen Liu, Mark Campbell, Kilian Q Weinberger, Bharath Hariharan, Wei-Lun Chao, Katie Z Luo

机构 * The Ohio State University(俄亥俄州立大学) Cornell University(康奈尔大学) Boston University(波士顿大学) Stanford University(斯坦福大学)

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 AutoReg3D通过自回归序列生成方法实现3D目标检测,无需锚点或NMS,展示了在LiDAR检测中的可行性与灵活性。

Comments CVPR 2026 Findings Project Page: https://tzmhuang.github.io/autoreg3d/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07464 2026-03-10 cs.CV 57%

Selective Transfer Learning of Cross-Modality Distillation for Monocular 3D Object Detection

跨模态蒸馏的 selective transfer learning 用于单目3D物体检测

Rui Ding, Meng Yang, Nanning Zheng

机构 * Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(人工智能与机器人研究所,西安交通大学)

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 本文提出MonoSTL方法,通过解决跨模态蒸馏中的模态差距问题,提升单目3D物体检测的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18853 2026-03-10 cs.CV 57%

Open-Vocabulary Domain Generalization in Urban-Scene Segmentation

面向城市场景分割的开放词汇领域泛化

Dong Zhao, Qi Zang, Nan Pu, Wenjing Li, Nicu Sebe, Zhun Zhong

机构 * Department of Information Engineering and Computer Science, University of Trento(信息工程与计算机科学系,特伦托大学) School of Computer Science and Information Engineering, Hefei University of Technology(计算机科学与信息工程学院,合肥工业大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 本文提出OVDG-SS,通过S2-Corr机制提升城市场景分割中开放词汇领域的泛化能力和鲁棒性。

Journal ref CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01487 2026-03-10 cs.CV 57%

PointSlice: Accurate and Efficient Slice-Based Representation for 3D Object Detection from Point Clouds

PointSlice: 用于点云3D物体检测的准确且高效的基于切片的表示方法

Liu Qifeng, Zhao Dawei, Dong Yabo, Xiao Liang, Wang Juan, Min Chen, Li Fuyang, Jiang Weizhong, Lu Dongming, Nie Yiming

机构 * Zhejiang University(浙江大学) Defense Innovation Institute(国防科技创新研究院) Tsinghua University(清华大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 PointSlice提出了一种基于切片的点云表示方法,通过引入切片交互网络提升3D物体检测的精度与效率。

Comments Accepted by Pattern Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02858 2026-03-10 cs.CV 57%

Empowering Microscopic Traffic Simulators with Realistic Perception using Surrogate Sensor Models

通过使用替代传感器模型使微观交通模拟器具备现实感知能力

Tianheng Zhu, Yiheng Feng

机构 * Lyles School of Civil and Construction Engineering, Purdue University(普渡大学莱尔斯土木与建设工程学院)

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 MIDAR通过替代传感器模型使微观交通模拟器具备现实感知能力,提升ITS应用的仿真真实性与效率。

Comments 27 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08438 2026-03-10 eess.SP 50%

Graph Based Semantic Encoder Decoder Framework for Task Oriented Communications in Connected Autonomous Vehicles

基于图的语义编码解码框架用于连接自动驾驶车辆的任务导向通信

Soheyb Ribouh, Phil Polo Ditsia Di Ngoma

专题命中 感知 :autonomous driving(abstract)

AI总结 本文提出基于图的语义编码解码框架,用于连接自动驾驶车辆的任务导向通信,通过语义压缩和解码提升通信效率与语义保真度。

详情

展开后加载摘要…

URL PDF HTML 收藏