arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

自动驾驶

自动驾驶感知、规划、BEV、占用预测、激光雷达和仿真评测。

共收录 6065 信号源:cs.RO, cs.CV, eess.IV, cs.AI

1. 感知 6065 篇

2503.13883 2026-01-21 cs.CV 57%

YOLO-LLTS: Real-Time Low-Light Traffic Sign Detection via Prior-Guided Enhancement and Multibranch Feature Interaction

YOLO-LLTS: 通过先验引导增强和多分支特征交互实现实时低光交通标志检测

Ziyu Lin, Yunfan Wu, Yuhang Ma, Junzhou Chen, Ronghui Zhang, Jiaming Wu, Guodong Yin, Liang Lin

机构 * Guangdong Key Laboratory of Intelligent Transportation System, School of intelligent systems engineering, Sun Yat-sen University(广东智能交通系统重点实验室,智能系统工程学院,中山大学) Department of Architecture and Civil Engineering, Chalmers University of Technology(建筑与土木工程系,查尔姆斯理工大学) School of Mechanical Engineering, Southeast University(机械工程学院,东南大学) School of Computer Science and Engineering, Sun Yat-sen University(计算机科学与工程学院,中山大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 YOLO-LLTS通过先验引导增强和多分支特征交互提升低光环境下交通标志检测精度。

Comments This work has been published in IEEE Transactions on Instrumentation and Measurement

Journal ref IEEE Trans. Instrum. Meas., vol. 74, pp. 1-18, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10819 2026-01-19 cs.CV 57%

A Unified 3D Object Perception Framework for Real-Time Outside-In Multi-Camera Systems

面向实时多相机系统的统一3D物体感知框架

Yizhou Wang, Sameer Pusegaonkar, Yuxing Wang, Anqi Li, Vishal Kumar, Chetan Sethi, Ganapathy Aiyer, Yun He, Kartikay Thakkar, Swapnil Rathi, Bhushan Rupde, Zheng Tang, Sujit Biswas

机构 * NVIDIA Corporation(英伟达公司)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 本文提出了一种面向实时多相机系统的统一3D物体感知框架,通过优化的Sparse4D框架和生成式数据增强策略,实现了在大规模基础设施环境中的高效多目标跟踪与实时部署。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08355 2026-01-16 cs.CV 57%

Semantic Misalignment in Vision-Language Models under Perceptual Degradation

视觉-语言模型在感知退化下的语义错位

Guo Cheng

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 本研究探讨了视觉-语言模型在感知退化下的语义错位问题,发现传统分割指标的下降并未影响下游行为,揭示了像素鲁棒性与多模态语义可靠性之间的脱节。

Comments 10 pages, 4 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22972 2026-01-16 cs.CV eess.SP 57%

Wavelet-based Multi-View Fusion of 4D Radar Tensor and Camera for Robust 3D Object Detection

基于小波的4D雷达张量与相机多视图融合用于鲁棒3D目标检测

Runwei Guan, Jianan Liu, Shaofeng Liang, Fangqiang Ding, Shanliang Yao, Xiaokai Bai, Daizong Liu, Tao Huang, Guoqiang Mao, Hui Xiong

机构 * Thrust of Artificial Intelligence, Hong Kong University of Science and Technology (Guangzhou)(人工智能 thrust,香港科技大学(广州)) Momoniai AI Department of Mechanical Engineering, Massachusetts Institute of Technology(机械工程系,麻省理工学院) School of Information Engineering, Yancheng Institute of Technology(信息工程学院,盐城科技学院) College of Information Science and Electronic Engineering, Zhejiang University(信息科学与电子工程学院,浙江大学) Institute for Math & AI, Wuhan University(数学与人工智能研究所,武汉大学) College of Science and Engineering and the Centre for AI and Data Science Innovation, James Cook University(科学与工程学院及人工智能与数据科学创新中心,詹姆斯库克大学) School of Transportation, Southeast University(交通运输学院,东南大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 WRCFormer通过小波注意力模块和几何引导渐进融合机制,高效融合4D雷达张量与相机图像,提升3D目标检测在恶劣天气下的鲁棒性。

Comments 10 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13601 2026-01-16 cs.CV 57%

Unleashing Semantic and Geometric Priors for 3D Scene Completion

释放语义和几何先验以实现3D场景补全

Shiyuan Chen, Wei Sui, Bohao Zhang, Zeyd Boukhers, John See, Cong Yang

机构 * D-Robotics

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 FoundationSSC通过双解耦机制和轴感知融合模块,提升3D场景补全的语义和几何指标表现。

Comments Accept by AAAI-2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13165 2026-01-15 quant-ph cs.AI cs.LG 57%

QuFeX: Quantum feature extraction module for hybrid quantum-classical deep neural networks

QuFeX:用于混合量子-经典深度神经网络的量子特征提取模块

Naman Jain, Amir Kalev

机构 * Viterbi School of Engineering, University of Southern California, Los Angeles, California 90089, USA Information Sciences Institute, University of Southern California, Arlington, VA 22203, USA Department of Physics Center for Quantum Information Science \& Technology, University of Southern California, Los Angeles, California 90089, USA

专题命中 感知 :autonomous driving(abstract);分类 cs.AI

AI总结 QuFeX通过在混合量子-经典深度神经网络中引入量子特征提取模块,提升了图像分割任务的性能。

Comments V2: 17 pages, 15 figures, 2 Tables; published version

Journal ref Quantum Sci. Technol. 11 015017 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08042 2026-01-14 physics.app-ph cs.RO 57%

μDopplerTag: CNN-Based Drone Recognition via Cooperative Micro-Doppler Tagging

μDopplerTag: 基于协作微多普勒标记的CNN无人机识别

O. Yerushalimov, D. Vovchuk, A. Glam, P. Ginzburg

专题命中 感知 :LiDAR(abstract);分类 cs.RO

AI总结 本文提出基于电磁标签和CNN的无人机识别方法,利用微多普勒签名实现远距离高精度分类,适用于空域监控等关键应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12735 2026-01-14 cs.CV 57%

Backdoor Attacks on Open Vocabulary Object Detectors via Multi-Modal Prompt Tuning

通过多模态提示调优对开放词汇目标检测器进行后门攻击

Ankita Raj, Chetan Arora

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 TrAP通过多模态提示调优对开放词汇目标检测器实施后门攻击,利用轻量级提示标记植入恶意行为,提升攻击成功率并改进下游任务性能。

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07168 2026-01-14 cs.CV 57%

HisTrackMap: Global Vectorized High-Definition Map Construction via History Map Tracking

HisTrackMap: 通过历史地图追踪构建全局向量高精度地图

Jing Yang, Sen Yang, Xiao Tan, Hanli Wang

机构 * Tongji University(同济大学) Baidu Inc.(百度公司)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 HisTrackMap通过历史地图追踪构建全局向量高精度地图,提升时间连续性和几何构造质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22200 2026-01-13 cs.RO 57%

EnvoDat: A Large-Scale Multisensory Dataset for Robotic Spatial Awareness and Semantic Reasoning in Heterogeneous Environments

EnvoDat:一种大规模多感官数据集,用于机器人空间感知和异构环境中的语义推理

Linus Nwankwo, Bjoern Ellensohn, Vedant Dave, Peter Hofer, Jan Forstner, Marlene Villneuve, Robert Galler, Elmar Rueckert

机构 * Chair of Cyber-Physical System, Montanuniversität Leoben, Austria(智能物理系统系,莱布恩矿业大学,奥地利) Theresianische Militarakademie, Austria(特里西亚军事学院,奥地利) Chair of Subsurface Engineering, Montanuniversität Leoben, Austria(地下工程系,莱布恩矿业大学,奥地利)

专题命中 感知 :autonomous driving(abstract);分类 cs.RO

AI总结 EnvoDat是一个大规模多感官数据集,用于提升机器人在复杂异构环境中的空间感知和语义推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19290 2026-01-13 cs.CV 57%

TRASE: Tracking-free 4D Segmentation and Editing

TRASE:无需跟踪的4D分割与编辑

Yun-Jin Li, Mariia Gladkova, Yan Xia, Daniel Cremers

机构 * TU Munich(慕尼黑技术大学) Munich Center for Machine Learning(慕尼黑机器学习中心)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 TRASE通过弱监督学习实现无需跟踪的4D分割,利用对比学习和聚类技术实现动态场景的高效分割与交互编辑。

Comments Accepted to 3DV 2026. Project page https://yunjinli.github.io/project-sadg

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04605 2026-01-09 cs.CV 57%

Detection of Deployment Operational Deviations for Safety and Security of AI-Enabled Human-Centric Cyber Physical Systems

面向AI赋能的人本型网络物理系统的部署操作偏差检测

Bernard Ngabonziza, Ayan Banerjee, Sandeep K. S. Gupta

专题命中 感知 :self-driving(abstract);分类 cs.CV

AI总结 本文提出了一种基于个性化图像的新技术,用于检测AI赋能的人本型网络物理系统在部署中的操作偏差,以确保其安全性和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06282 2026-01-09 cs.CV 57%

From Dataset to Real-world: General 3D Object Detection via Generalized Cross-domain Few-shot Learning

从数据集到现实世界:通过通用跨领域少样本学习实现通用3D目标检测

Shuangzhi Li, Junlong Shen, Lei Ma, Xingyu Li

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 本文提出通用跨领域少样本学习方法,通过融合2D语义与3D空间推理,实现对现实世界中常见和新类目标的高效检测。

Comments The latest version refines the few-shot setting on common classes, enforcing a stricter object-level definition

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03301 2026-01-08 cs.MA cs.AI 57%

PC2P: Multi-Agent Path Finding via Personalized-Enhanced Communication and Crowd Perception

PC2P:通过个性化增强通信与人群感知进行多智能体路径寻找

Guotao Li, Shaoyun Xu, Yuexing Hao, Yang Wang, Yuhui Sun

机构 * Institute of Microelectronics of the Chinese Academy of Sciences(中国科学院微电子研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 感知 :occupancy(abstract);分类 cs.AI

AI总结 PC2P通过个性化增强通信与人群感知方法,提升多智能体路径寻找在复杂环境中的协同与扩展能力。

Comments 8 pages,7 figures,3 tables,Accepted to IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01695 2026-01-06 cs.CV 57%

Learnability-Driven Submodular Optimization for Active Roadside 3D Detection

基于可学习性的子模优化用于主动道路3D检测

Ruiyu Mao, Baoming Zhang, Nicholas Ruozzi, Yunhui Guo

机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 本研究提出了一种基于可学习性的主动学习框架,用于道路侧单目3D物体检测,通过选择信息丰富且可可靠标注的场景,有效减少标注成本并提升模型性能。

Comments 10 pages, 7 figures. Submitted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01676 2026-01-06 cs.CV 57%

LabelAny3D: Label Any Object 3D in the Wild

LabelAny3D: 在真实世界中标注任意物体的3D

Jin Yao, Radowan Mahmud Redoy, Sebastian Elbaum, Matthew B. Dwyer, Zezhou Cheng

机构 * University of Virginia(弗吉尼亚大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 LabelAny3D通过分析-合成框架生成高质量3D标注,提升单目3D检测性能,推动真实世界3D识别发展。

Comments NeurIPS 2025. Project page: https://uva-computer-vision-lab.github.io/LabelAny3D/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00368 2026-01-05 cs.CV 57%

Mask-Conditioned Voxel Diffusion for Joint Geometry and Color Inpainting

基于掩码的体素扩散用于联合几何和颜色修复

Aarya Sumuk

专题命中 感知 :occupancy(abstract);分类 cs.CV

AI总结 本文提出基于掩码的体素扩散方法,用于修复受损3D物体的几何和颜色,通过两阶段框架实现更完整和一致的修复效果。

Comments 10 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24680 2026-01-01 cs.RO 57%

ReSPIRe: Informative and Reusable Belief Tree Search for Robot Probabilistic Search and Tracking in Unknown Environments

ReSPIRe: 信息性和可重用的信念树搜索用于未知环境中的机器人概率搜索与跟踪

Kangjie Zhou, Zhaoyang Li, Han Gao, Yao Su, Hangxin Liu, Junzhi Yu, Chang Liu

专题命中 感知 :trajectory planning(abstract);分类 cs.RO

AI总结 ReSPIRe通过分层粒子结构和可重用信念树搜索,在未知环境中实现高效且稳定的机器人目标搜索与跟踪。

Comments 17 pages, 12 figures, accepted to IEEE Transactions on Systems, Man, and Cybernetics: Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24243 2026-01-01 cs.CV 57%

MambaSeg: Harnessing Mamba for Accurate and Efficient Image-Event Semantic Segmentation

MambaSeg: 利用Mamba实现准确且高效的图像事件语义分割

Fuqiang Gu, Yuanke Li, Xianlei Long, Kangping Ji, Chao Chen, Qingyi Gu, Zhenliang Ni

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 MambaSeg通过双分支框架和双维交互模块,实现高效准确的多模态语义分割,适用于快速运动和低光条件下的图像事件分割任务。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24227 2026-01-01 cs.CV 57%

Mirage: One-Step Video Diffusion for Photorealistic and Coherent Asset Editing in Driving Scenes

Mirage: 驾驶场景中光实且一致的资产编辑一步视频扩散

Shuyun Wang, Haiyang Sun, Bing Wang, Hangjun Ye, Xin Yu

机构 * The University of Queensland(昆士兰大学) Xiaomi EV(小米电动车)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 Mirage提出了一种用于驾驶场景中光实且一致的资产编辑的一步视频扩散模型,通过引入时序无关的潜在特征和两阶段数据对齐策略,提升了视觉保真度和时间一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23635 2025-12-30 cs.CV 57%

Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception

重新思考端到端3D感知的时空对齐

Xiaoyu Li, Peidong Li, Xian Wu, Long Shi, Dedong Liu, Yitao Wu, Jiajia Fu, Dixiao Cui, Lijun Zhao, Lining Sun

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 HAT通过自适应解码多假设生成最优时空对齐提案,提升自动驾驶中3D感知精度和鲁棒性。

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23585 2025-12-30 cs.RO cs.SY eess.SY 57%

Unsupervised Learning for Detection of Rare Driving Scenarios

无监督学习用于罕见驾驶场景检测

Dat Le, Thomas Manhardt, Moritz Venator, Johannes Betz

机构 * Professorship of Autonomous Vehicle Systems, TUM School of Engineering and Design, Technical University Munich(自主车辆系统教授职位,技术大学慕尼黑工程与设计学院) Munich Institute of Robotics and Machine Intelligence (MIRMI)(慕尼黑机器人与机器智能研究所) CARIAD SE(CARIAD公司)

专题命中 感知 :autonomous driving(abstract);分类 cs.RO

AI总结 本研究提出一种无监督学习方法,利用深度隔离森林和t-SNE技术,有效检测自动驾驶中的罕见危险场景。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22513 2025-12-30 eess.SP eess.IV 57%

CoDS: Collaborative Perception via Digital Semantic Communication

CoDS:基于数字语义通信的协同感知

Jipeng Gan, Le Liang, Hua Zhang, Chongtao Guo, Shi Jin

专题命中 感知 :autonomous driving(abstract);分类 eess.IV

AI总结 CoDS通过数字语义通信实现协同感知,解决传统方法在V2X网络中的兼容性问题,提升感知性能与传输效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22392 2025-12-30 cs.CV 57%

iOSPointMapper: RealTime Pedestrian and Accessibility Mapping with Mobile AI

iOSPointMapper: 基于移动AI的实时行人及可达性制图

Himanshu Naidu, Yuxiang Zhang, Sachin Mehta, Anat Caspi

机构 * University of Washington(华盛顿大学) Paul G. Allen School of Computer Science and Engineering(保罗·G·艾伦计算机科学与工程学院)

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 iOSPointMapper通过移动AI实现实时人行道制图,结合语义分割、LiDAR和GPS数据,提供隐私保护的数据收集与多模式交通数据整合,提升行人基础设施的可达性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20815 2025-12-29 cs.CV 57%

Learning to Sense for Driving: Joint Optics-Sensor-Model Co-Design for Semantic Segmentation

学习感知以驾驶:语义分割的联合光学-传感器-模型协同设计

Reeshad Khan, John Gauch

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 本文提出了一种联合光学-传感器-模型协同设计框架,通过端到端RAW到任务流水线提升语义分割性能,实现在边缘设备上的高效部署。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19083 2025-12-25 cs.RO 57%

CoDrone: Autonomous Drone Navigation Assisted by Edge and Cloud Foundation Models

CoDrone:由边缘和云基础模型辅助的自主无人机导航

Pengyu Chen, Tao Ouyang, Ke Luo, Weijie Hong, Xu Chen

专题命中 感知 :occupancy(abstract);分类 cs.RO

AI总结 CoDrone通过整合云-边-端协作计算框架和基础模型,提升无人机自主导航性能,实现更高效和精确的环境感知与动态适应。

Comments This paper is accepted by the IEEE Internet of Things Journal (IoT-J) for publication in the Special Issue on "Augmented Edge Sensing Intelligence for Low-Altitude IoT Systems"

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20217 2025-12-24 cs.CV 57%

LiteFusion: Taming 3D Object Detectors from Vision-Based to Multi-Modal with Minimal Adaptation

LiteFusion: 从基于视觉到多模态的3D目标检测器的最小适应

Xiangxuan Ren, Zhongdao Wang, Pin Tang, Guoqing Wang, Jilai Zheng, Chao Ma

机构 * China Ministry of Education (MOE) Key Laboratory of Artificial Intelligence(中国教育部人工智能重点实验室) Artificial Intelligence Institute, Shanghai Jiao Tong University(上海交通大学人工智能学院) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 LiteFusion通过将LiDAR数据作为补充几何信息源,无需专用LiDAR编码器,显著提升多模态3D目标检测性能。

Comments 13 pages, 9 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15109 2025-12-23 eess.SP cs.AI cs.IT math.IT 57%

Large Model Enabled Embodied Intelligence for 6G Integrated Perception, Communication, and Computation Network

大模型赋能的具身智能用于6G集成感知、通信与计算网络

Zhuoran Li, Zhen Gao, Xinhua Liu, Zheng Wang, Xiaotian Zhou, Lei Liu, Yongpeng Wu, Wei Feng, Yongming Huang

机构 * School of Information and Electronics(信息与电子学院) School of Interdisciplinary Science(交叉科学学院) State Key Laboratory of CNS/ATM(CNS/ATM国家重点实验室) Beijing Institute of Technology(北京理工大学) Advanced Technology Research Institute(先进技术研究院) Yangtze Delta Region Academy(长江三角洲地区学院) School of School of Information Science and Engineering(信息科学与工程学院) Institute of Intelligent Communication Technologies(智能通信技术研究院) Shandong Key Laboratory of Intelligent Communication and Sensing-Computing Integration(智能通信与传感-计算集成山东省重点实验室) Zhejiang Provincial Key Laboratory of Information Processing, Communication and Networking(信息处理、通信与网络浙江省重点实验室) Department of Electronic Engineering(电子工程系) State Key Laboratory of Space Network and Communications(空间网络与通信国家重点实验室) National Mobile Communications Research Laboratory(移动通信研究中心) Purple Mountain Laboratories(紫金山实验室)

专题命中 感知 :autonomous driving(abstract);分类 cs.AI

AI总结 本文提出利用大模型赋能基站实现感知、通信和计算一体化,为6G系统提供安全关键的智能解决方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13982 2025-12-23 cs.CV 57%

FocalComm: Hard Instance-Aware Multi-Agent Perception

FocalComm: 重视困难实例的多智能体感知

Dereje Shenkut, Vijayakumar Bhagavatula

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 FocalComm通过聚焦困难实例特征交换,提升多智能体协作感知中对行人等安全关键物体的检测性能。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17450 2025-12-22 cs.CV cs.LG 57%

MULTIAQUA: A multimodal maritime dataset and robust training strategies for multimodal semantic segmentation

MULTIAQUA:一种多模态海洋数据集和多模态语义分割的鲁棒训练策略

Jon Muhovič, Janez Perš

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 MULTIAQUA数据集和多模态语义分割的鲁棒训练策略,旨在通过多模态数据提升恶劣能见度下的场景解释质量。

详情

展开后加载摘要…

URL PDF HTML 收藏