arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

自动驾驶

自动驾驶感知、规划、BEV、占用预测、激光雷达和仿真评测。

共收录 6065 信号源:cs.RO, cs.CV, eess.IV, cs.AI

1. 感知 6065 篇

2308.05818 2026-03-12 cs.CV eess.SP 57%

Absorption-Based, Passive Range Imaging from Hyperspectral Thermal Measurements

基于吸收的被动超光谱热测量范围成像

Unay Dorken Gallastegi, Hoover Rueda-Chacon, Martin J. Stevens, Vivek K Goyal

机构 * Department of Electrical and Computer Engineering, Boston University(波士顿大学电气与计算机工程系) Department of Computer Science, Universidad Industrial de Santander(圣安德烈斯工业大学计算机科学系) National Institute of Standards and Technology(美国国家标准与技术研究院)

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 本文提出一种基于超光谱热测量的被动范围成像方法,通过计算分离热辐射和大气吸收效应,实现无主动照明下的远距离物体范围估计。

Comments 15 pages, 14 figures

Journal ref IEEE Trans. Pattern Analysis & Machine Intelligence, vol. 47, no. 5, pp. 4044-4060, May 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.04796 2026-03-12 cs.CV 57%

In Pursuit of Many: A Review of Modern Multiple Object Tracking Systems

追寻众多:现代多目标跟踪系统的综述

Mk Bashar, Samia Islam, Kashifa Kawaakib Hussain, Md. Bakhtiar Hasan, A. B. M. Ashikur Rahman, Md. Hasanul Kabir

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 本文综述了现代多目标跟踪系统的发展,涵盖从基于检测到端到端设计的演变,以及基于变压器、生成模型等架构的最新进展,并探讨了基准趋势和实际部署中的挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07737 2026-03-11 cs.CV 57%

SpikeSMOKE: Spiking Neural Networks for Monocular 3D Object Detection with Cross-Scale Gated Coding

SpikeSMOKE:用于单目3D目标检测的脉冲神经网络与跨尺度门控编码

Xuemei Chen, Huamin Wang, Jing Peng, Hangchi Shen, Shukai Duan, Shiping Wen, Tingwen Huang

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 SpikeSMOKE通过跨尺度门控编码机制提升单目3D目标检测的低功耗性能,实现参数和计算量的显著减少

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08379 2026-03-10 cs.RO 57%

Perception-Aware Communication-Free Multi-UAV Coordination in the Wild

野外多无人机协同的感知-aware 通信-free 方法

Manuel Boldrer, Michal Kamler, Afzal Ahmad, Martin Saska

机构 * Department of Cybernetics, Czech Technical University in Prague(控制系,捷克技术大学)

专题命中 感知 :LiDAR(abstract);分类 cs.RO

AI总结 本研究提出了一种无需通信的多无人机协同方法,利用3D激光雷达进行SLAM和障碍物检测,实现复杂环境中的安全导航。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08240 2026-03-10 cs.CV 57%

SiMO: Single-Modality-Operable Multimodal Collaborative Perception

SiMO: 单模态可操作的多模态协作感知

Jiageng Wen, Shengjie Zhao, Bing Li, Jiafeng Huang, Kenan Ye, Hao Deng

机构 * Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University(同济大学智能自主系统上海研究院) School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院) School of Mechatronic Engineering and Automation, Shanghai University(上海大学机械电子工程与自动化学院)

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 SiMO通过单模态可操作的多模态协作感知方法,解决多模态特征融合中的语义不匹配问题,提升协作感知性能。

Comments Accepted to ICLR 2026. This arXiv version includes an additional appendix (Appendix 15) containing further philosophical discussion not included in the official ICLR peer-reviewed version

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00661 2026-03-10 cs.CV 57%

Elytra: A Flexible Framework for Securing Large Vision Systems

ELYTRA:一种用于安全大型视觉系统的灵活框架

Richard E. Neddo, Emmanuel Atindama, Zander W. Blasingame, Chen Liu

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 ELYTRA框架通过低秩适应技术,为大型视觉系统提供动态安全修补,有效提升对抗性示例下的分类准确率

Comments Updated pre-print. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07985 2026-03-10 cs.CV 57%

On the Feasibility and Opportunity of Autoregressive 3D Object Detection

关于自回归3D目标检测的可行性与机会

Zanming Huang, Jinsu Yoo, Sooyoung Jeon, Zhenzhen Liu, Mark Campbell, Kilian Q Weinberger, Bharath Hariharan, Wei-Lun Chao, Katie Z Luo

机构 * The Ohio State University(俄亥俄州立大学) Cornell University(康奈尔大学) Boston University(波士顿大学) Stanford University(斯坦福大学)

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 AutoReg3D通过自回归序列生成方法实现3D目标检测,无需锚点或NMS,展示了在LiDAR检测中的可行性与灵活性。

Comments CVPR 2026 Findings Project Page: https://tzmhuang.github.io/autoreg3d/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07464 2026-03-10 cs.CV 57%

Selective Transfer Learning of Cross-Modality Distillation for Monocular 3D Object Detection

跨模态蒸馏的 selective transfer learning 用于单目3D物体检测

Rui Ding, Meng Yang, Nanning Zheng

机构 * Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(人工智能与机器人研究所,西安交通大学)

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 本文提出MonoSTL方法,通过解决跨模态蒸馏中的模态差距问题,提升单目3D物体检测的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18853 2026-03-10 cs.CV 57%

Open-Vocabulary Domain Generalization in Urban-Scene Segmentation

面向城市场景分割的开放词汇领域泛化

Dong Zhao, Qi Zang, Nan Pu, Wenjing Li, Nicu Sebe, Zhun Zhong

机构 * Department of Information Engineering and Computer Science, University of Trento(信息工程与计算机科学系,特伦托大学) School of Computer Science and Information Engineering, Hefei University of Technology(计算机科学与信息工程学院,合肥工业大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 本文提出OVDG-SS,通过S2-Corr机制提升城市场景分割中开放词汇领域的泛化能力和鲁棒性。

Journal ref CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01487 2026-03-10 cs.CV 57%

PointSlice: Accurate and Efficient Slice-Based Representation for 3D Object Detection from Point Clouds

PointSlice: 用于点云3D物体检测的准确且高效的基于切片的表示方法

Liu Qifeng, Zhao Dawei, Dong Yabo, Xiao Liang, Wang Juan, Min Chen, Li Fuyang, Jiang Weizhong, Lu Dongming, Nie Yiming

机构 * Zhejiang University(浙江大学) Defense Innovation Institute(国防科技创新研究院) Tsinghua University(清华大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 PointSlice提出了一种基于切片的点云表示方法,通过引入切片交互网络提升3D物体检测的精度与效率。

Comments Accepted by Pattern Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02858 2026-03-10 cs.CV 57%

Empowering Microscopic Traffic Simulators with Realistic Perception using Surrogate Sensor Models

通过使用替代传感器模型使微观交通模拟器具备现实感知能力

Tianheng Zhu, Yiheng Feng

机构 * Lyles School of Civil and Construction Engineering, Purdue University(普渡大学莱尔斯土木与建设工程学院)

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 MIDAR通过替代传感器模型使微观交通模拟器具备现实感知能力,提升ITS应用的仿真真实性与效率。

Comments 27 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05355 2026-03-09 cs.RO 57%

OmniDP: Beyond-FOV Large-Workspace Humanoid Manipulation with Omnidirectional 3D Perception

OmniDP: 在超广角大工作空间中实现人形机器人的全方位3D感知

Pei Qu, Zheng Li, Yufei Jia, Ziyun Liu, Liang Zhu, Haoang Li, Jinni Zhou, Jun Ma

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Tsinghua University(清华大学)

专题命中 感知 :LiDAR(abstract);分类 cs.RO

AI总结 OmniDP通过360度全景点云感知和时间感知注意力池化机制,实现人形机器人在大工作空间中的稳健操作,优于传统深度相机方法。

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09173 2026-03-09 cs.CV 57%

Beyond Flat Unknown Labels in Open-World Object Detection

超越开放世界中模糊未知标签的物体检测

Yuchen Zhang, Yao Lu, Johannes Betz

机构 * AVS Lab, Technical University of Munich(慕尼黑技术大学avs实验室)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 本文提出BOUND,一种开放世界物体检测器,通过推断未知物体的粗粒度类别提升检测性能,同时实现结构化的层次分类。

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14266 2026-03-06 cs.CV 57%

DriverGaze360: OmniDirectional Driver Attention with Object-Level Guidance

DriverGaze360: 全向驾驶员注意力与物体级引导

Shreedhar Govil, Didier Stricker, Jason Rambach

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 DriverGaze360通过全向视野和物体级引导方法,实现对驾驶员注意力的高精度预测,提升自动驾驶系统中的空间感知与交互能力。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04470 2026-03-06 cs.RO 57%

Efficient Autonomous Navigation of a Quadruped Robot in Underground Mines on Edge Hardware

在边缘硬件上实现四足机器人在地下矿井中高效自主导航

Yixiang Gao, Kwame Awuah-Offei

机构 * Department of Mining and Explosive Engineering, Missouri University of Science and Technology(采矿与爆炸工程系,密苏里科技大学)

专题命中 感知 :LiDAR(abstract);分类 cs.RO

AI总结 本研究提出了一种在边缘硬件上运行的四足机器人自主导航系统,无需GPU和网络连接,实现了在地下矿井中高效、可靠的自主导航。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12567 2026-03-04 cs.CV 57%

GAN-Based Single-Stage Defense for Traffic Sign Classification Under Adversarial Patch

基于GAN的单阶段交通标志分类对抗攻击防御策略

Abyad Enan, Mashrur Chowdhury

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 本文提出基于GAN的单阶段防御策略,有效提升自动驾驶系统对对抗性贴纸攻击的防御能力,显著提高交通标志分类的准确率。

Comments This work has been submitted to a peer-reviewed journal and is currently under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02560 2026-03-04 cs.CV 57%

CAWM-Mamba: A unified model for infrared-visible image fusion and compound adverse weather restoration

CAWM-Mamba:一种用于红外可见图像融合和复合恶劣天气恢复的统一模型

Huichun Liu, Xiaosong Li, Zhuangfan Huang, Tao Ye, Yang Liu, Haishu Tan

机构 * School of Physics(物理学院) Optoelectronic Engineering, Foshan University, Foshan 528225, China(光学电子工程学院,佛山大学,佛山528225,中国) Guangdong-HongKong-Macao Joint Laboratory for Intelligent Micro-Nano Optoelectronic Technology, Foshan 528225, China(粤港澳联合智能微纳光电子技术实验室,佛山528225,中国) School of Mechanical Electronic(机械电子学院) Information Engineering, China University of Mining and Technology, Beijing 100083, China(信息工程学院,中国矿业大学,北京100083,中国)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 CAWM-Mamba提出了一种统一模型,用于红外可见图像融合和复合恶劣天气恢复,通过三个关键模块提升多退化场景下的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02532 2026-03-04 cs.CV 57%

EIMC: Efficient Instance-aware Multi-modal Collaborative Perception

EIMC: 高效实例感知多模态协作感知

Kang Yang, Peng Wang, Lantao Li, Tianci Bu, Chen Sun, Deying Li, Yongcai Wang

机构 * School of Information, Renmin University of China(中国人民大学信息学院) Sony Research and Development Center China(索尼(中国)研发有限公司) National University of Defense Technology(国防科技大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 EIMC通过实例感知的多模态协作感知方法,提升自动驾驶安全性,减少带宽使用,实现高效且准确的3D感知。

Comments 9 pages, 8 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11772 2026-03-03 cs.CV 57%

Seg2Track-SAM2: SAM2-based Multi-object Tracking and Segmentation

Seg2Track-SAM2:基于SAM2的多目标跟踪与分割

Diogo Mendonça, Tiago Barros, Cristiano Premebida, Urbano J. Nunes

机构 * University of Coimbra(科英布拉大学) Institute of Systems and Robotics(系统与机器人研究所) Department of Electrical and Computer Engineering(电气与计算机工程系)

专题命中 感知 :driving perception(abstract);分类 cs.CV

AI总结 Seg2Track-SAM2基于SAM2提出多目标跟踪与分割框架,通过整合预训练检测器与专用模块,实现跟踪初始化、数据关联和细化,提升身份一致性和内存效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01847 2026-03-03 cs.CV 57%

GroupEnsemble: Efficient Uncertainty Estimation for DETR-based Object Detection

GroupEnsemble: 用于基于DETR的目标检测的高效不确定性估计

Yutong Yang, Katarina Popović, Julian Wiederer, Markus Braun, Vasileios Belagiannis, Bin Yang

机构 * Mercedes-Benz AG(梅赛德斯-奔驰集团) University of Stuttgart(斯图加特大学) Friedrich-Alexander University Erlangen-Nuremberg(埃尔兰根-纽伦堡弗里德里希-亚历山大大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 GroupEnsemble通过高效方法提升DETR模型的空间不确定性估计,结合MC-Dropout在自动驾驶场景中实现更高效的检测可靠性评估。

Comments Accepted to IEEE IV 2026. 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01708 2026-03-03 cs.CV 57%

WhisperNet: A Scalable Solution for Bandwidth-Efficient Collaboration

WhisperNet: 一种可扩展的带宽高效协作解决方案

Gong Chen, Chaokun Zhang, Xinyan Zhao

机构 * School of Computer Science and Technology, Tianjin University(计算机科学与技术学院,天津大学) School of Cybersecurity, Tianjin University(网络安全学院,天津大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 WhisperNet通过引入以接收器为中心的全局协调框架,实现了带宽高效协作,提升协同感知性能。

Comments Accepted by CVPR26

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00338 2026-03-03 cs.RO 57%

Layered Safety: Enhancing Autonomous Collision Avoidance via Multistage CBF Safety Filters

分层安全:通过多阶段CBF安全过滤器增强自主避障

Erina Yamaguchi, Ryan M. Bena, Gilbert Bahati, Aaron D. Ames

机构 * Caltech(卡内基梅隆大学)

专题命中 感知 :occupancy(abstract);分类 cs.RO

AI总结 本文提出一种多阶段CBF安全过滤器,通过预测和实时安全过滤提升机器人动态避障的鲁棒性和安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23926 2026-03-03 cs.CV 57%

Point-MoE: Large-Scale Multi-Dataset Training with Mixture-of-Experts for 3D Semantic Segmentation

点-MoE:通过混合专家进行大规模多数据集训练用于3D语义分割

Xuweiyi Chen, Wentao Zhou, Aruni RoyChowdhury, Zezhou Cheng

机构 * University of Virginia(弗吉尼亚大学) MathWorks

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 Point-MoE通过混合专家方法在无需数据集标签的情况下,实现大规模多数据集联合训练,提升3D语义分割性能。

Comments Project page: https://point-moe.cs.virginia.edu/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22920 2026-02-27 cs.CV 57%

OSDaR-AR: Enhancing Railway Perception Datasets via Multi-modal Augmented Reality

OSDaR-AR: 通过多模态增强现实提升铁路感知数据集

Federico Nesti, Gianluca D'Amico, Mauro Marinoni, Giorgio Buttazzo

机构 * Department of Excellence in Robotics & AI, Scuola Superiore Sant’Anna(机器人与人工智能卓越部门,圣安娜高等学院) Simulatrix MV srl(Simulatrix MV公司)

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 本文提出OSDaR-AR数据集,通过多模态增强现实技术整合逼真虚拟对象,提升铁路感知任务的数据质量与真实性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02686 2026-02-27 cs.CV 57%

ClimaOoD: Improving Anomaly Segmentation via Physically Realistic Synthetic Data

ClimaOoD: 通过物理真实合成数据提升异常分割

Yuxing Liu, Zheng Li, Huanhuan Liang, Ji Zhang, Zeyu Sun, Yong Liu

机构 * Beijing University of Chemical Technology(北京化工大学) Southwest Minzu University(西南民族大学) Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 ClimaOoD通过生成物理真实且语义连贯的合成数据,提升异常分割的鲁棒性和泛化能力。

Comments Accepted by CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13587 2026-02-27 cs.CV 57%

UniFuture: A 4D Driving World Model for Future Generation and Perception

UniFuture: 一个面向未来生成与感知的4D驾驶世界模型

Dingkang Liang, Dingyuan Zhang, Xin Zhou, Sifan Tu, Tianrui Feng, Xiaofan Li, Yumeng Zhang, Mingyang Du, Xiao Tan, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) Baidu Inc.(百度公司)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 UniFuture通过统一4D建模提升自动驾驶中的未来生成与几何感知能力。

Comments Accepted by ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21992 2026-02-26 cs.CV 57%

PanoEnv: Exploring 3D Spatial Intelligence in Panoramic Environments with Reinforcement Learning

PanoEnv:探索基于强化学习的全景环境中3D空间智能

Zekai Lin, Xu Zheng

机构 * University of Glasgow(格拉斯哥大学) HKUST(GZ)(香港科技大学(广州))

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 PanoEnv通过强化学习框架提升VLMs在全景环境中3D空间推理能力,实现更高准确率和语义评分。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19994 2026-02-24 cs.CV 57%

RADE-Net: Robust Attention Network for Radar-Only Object Detection in Adverse Weather

RADE-Net:面向恶劣天气的雷达-only目标检测稳健注意力网络

Christof Leitgeb, Thomas Puchleitner, Max Peter Ronecker, Daniel Watzenig

机构 * Infineon Technologies AG(英飞凌科技有限公司) Institute of Visual Computing(视觉计算研究所) Graz University of Technology(格拉茨工业大学) Virtual Vehicle Research GmbH(虚拟车辆研究有限公司)

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 RADE-Net通过3D投影方法和轻量模型提升雷达在恶劣天气下的目标检测性能。

Comments Accepted to 2026 IEEE Intelligent Vehicles Symposium (IV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14363 2026-02-17 cs.RO cs.LG 57%

AdaptManip: Learning Adaptive Whole-Body Object Lifting and Delivery with Online Recurrent State Estimation

AdaptManip: 基于在线递归状态估计的自适应全身体态物体抓取与配送

Morgan Byrd, Donghoon Baek, Kartik Garg, Hyunyoung Jung, Daesol Cho, Maks Sorokin, Robert Wright, Sehoon Ha

专题命中 感知 :LiDAR(abstract);分类 cs.RO

AI总结 AdaptManip通过强化学习实现人形机器人自主导航、物体抓取与配送,无需人类示范,具备高鲁棒性和适应性。

Comments Website: https://morganbyrd03.github.io/adaptmanip/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13003 2026-02-16 cs.CV cs.LG 57%

MASAR: Motion-Appearance Synergy Refinement for Joint Detection and Trajectory Forecasting

MASAR: 运动-外观协同细化用于联合检测与轨迹预测

Mohammed Amine Bencheikh Lehocine, Julian Schmidt, Frank Moosmann, Dikshant Gupta, Fabian Flohr

机构 * Mercedes-Benz AG(梅赛德斯-奔驰集团) Munich University of Applied Sciences(慕尼黑应用科学大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 MASAR通过联合编码外观和运动特征,提升自动驾驶中的3D检测与轨迹预测性能,实现超过20%的精度提升。

Comments Accepted to the 2026 IEEE International Conference on Robotics and Automation (ICRA 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏