arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

自动驾驶

自动驾驶感知、规划、BEV、占用预测、激光雷达和仿真评测。

2026-08-04 至 2026-08-04 共收录 13 信号源:cs.RO, cs.CV, eess.IV, cs.AI

1. 感知 3 篇

2605.30239 2026-08-04 cs.CV 版本更新 57%

CA-World: Multi-Object Counterfactual Alignment for Efficient Interactive-Ready Reconstruction

SAM3D-Phys:迈向真实世界中的多物体交互仿真

Xin Dong, Weijian Deng, Lihan Zhang, Tianru Dai, Wenfeng Deng, Yansong Tang

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Pengcheng Laboratory(鹏城实验室)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 提出SAM3D-Phys框架,结合场景重建与SAM3D生成式先验,从部分观测中恢复完整可仿真物体几何,并通过物理约束优化和掩码引导外观蒸馏实现场景一致性,支持多物体同时交互仿真。

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29471 2026-08-04 cs.CV 版本更新 57%

V2VCrafter: Consistent Street-View Image Generation Across Vehicles

V2XCrafter:学习生成跨智能体的驾驶场景

Yihang Tao, Yu Guo, Senkang Hu, Yanan Ma, Zihan Fang, Sam Kwong, Yuguang Fang

机构 * Hong Kong JC STEM Lab of Smart City(香港JC智能城市STEM实验室) City University of Hong Kong(香港城市大学) Lingnan University(岭南大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 提出V2XCrafter框架,通过渐进式多智能体扩散模型和跨智能体注意力模块,生成跨智能体相机视角的一致可控协作驾驶场景,以增强数据并提升下游协作3D目标检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08163 2026-08-04 cs.CV 版本更新 57%

Accuracy Does Not Guarantee Human-Likeness: Cross-Domain Human-Centered Benchmark in Monocular Depth Estimation

单目深度估计器中准确性并不保证人类化

Yuki Kubota, Taiki Fukiage

机构 * Communication Science Laboratories, NTT, Inc.(NTT通信科学实验室)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 研究发现单目深度估计器的准确性与人类相似性之间存在复杂权衡,提升准确性不一定增强人类化行为。

Comments 21 pages, 14 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 规划控制 7 篇

2512.15038 2026-08-04 cs.AI 版本更新 83%

LADY: Linear Attention for Autonomous Driving Efficiency without Transformers

LADY:用于自动驾驶效率的线性注意力,无需Transformer

Jihao Huang, Xi Xia, Zhiyuan Li, Tianle Liu, Jingke Wang, Junbo Chen, Tengju Ye

机构 * Udeer AI Zhejiang University(浙江大学) Yuanshi Intelligence(元世智能)

专题命中 规划控制 :autonomous driving(title,abstract);LiDAR(abstract_cn);分类 cs.AI

AI总结 LADY是首个基于线性注意力的端到端自动驾驶生成模型,通过常数时间与内存复杂度实现高效长时序建模与跨模态交互。

Comments Accepted by IEEE RAL

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00919 2026-08-04 cs.CV cs.RO 版本更新 81%

DriveCode: Domain Specific Numerical Encoding for LLM-Based Autonomous Driving

DriveCode: 针对基于LLM的自动驾驶的领域特定数值编码

Zhiye Wang, Yanbo Jiang, Rui Zhou, Bo Zhang, Fang Zhang, Zhenhua Xu, Yaqin Zhang, Jianqiang Wang

机构 * School of Information Science and Engineering, Lanzhou University(兰州大学信息科学与工程学院) The School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动性学院) The Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院) DiDi, Beijing, China(滴滴出行) State Key Laboratory of Intelligent Green Vehicle and Mobility, Tsinghua University(清华大学智能绿色车辆与移动性国家重点实验室)

专题命中 规划控制 :autonomous driving(title,abstract);分类 cs.RO、cs.CV

AI总结 本文提出DriveCode,一种将数字表示为专用嵌入而非离散文本标记的新型数值编码方法,提升LLM在自动驾驶中的数值推理与解码效率。

Comments The project page is available at https://shiftwilliam.github.io/DriveCode

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15829 2026-08-04 cs.ET cs.AI cs.CV cs.NE 版本更新 62%

Human-like working memory signatures emerge from intrinsically plastic artificial neurons for robust dynamic vision

物理驱动的人类样工作记忆在动态视觉中优于数字网络

Jingli Liu, Huannan Zheng, Bohao Zou, Kezhou Yang

机构 * Microelectronics Thrust, The Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州)微电子研究组)

专题命中 规划控制 :autonomous driving(abstract);分类 cs.CV、cs.AI

AI总结 本文提出基于物理的Intrinsic Plasticity Network (IPNet),通过磁隧道结的焦耳加热弛豫动态实现类人工作记忆,显著提升动态视觉任务的效率与精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.19719 2026-08-04 cs.LG cs.RO 版本更新 57%

Koopman Dreamer: Spectrally Constrained Latent Dynamics for Stable World-Model Imagination

库普曼梦想家:用于稳定世界模型想象的频谱约束潜在动力学

Jiaqi Li, Xinglong Zhang, Haibin Xie, Yixing Lan, Wei Pan, Xin Xu

机构 * College of Intelligence Science and Technology, National University of Defense Technology(国防科技大学智能科学与技术学院) School of Engineering, Newcastle University(纽卡斯尔大学工程学院)

专题命中 规划控制 :LiDAR(abstract);分类 cs.RO

AI总结 研究针对潜在世界模型长期展开中模态持久性和误差积累控制有限的问题,提出库普曼梦想家模型,通过频谱约束潜在动力学核心及多种目标结合优化,推导误差界,实验证明其提高了长期潜在展开稳定性及闭环控制性能。

Comments 20 pages, 13 figures, 11 tables. Revised manuscript with a more concise and precise abstract and improved clarity and presentation throughout the main text. The main technical content, experimental results, and conclusions remain unchanged

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00500 2026-08-04 cs.AI cs.SE 版本更新 57%

ProbGuard: Proactive Runtime Monitoring for LLM Agent Safety via Probabilistic Prediction

ProbGuard:基于概率的运行时监控用于LLM代理安全

Haoyu Wang, Christopher M. Poskitt, Jiali Wei, Jun Sun

机构 * Singapore Management University(新加坡管理大学) Xi'an Jiaotong University(西安交通大学)

专题命中 规划控制 :autonomous driving(abstract);分类 cs.AI

AI总结 ProbGuard通过概率风险预测主动监控LLM代理安全,利用离散时间马尔可夫链建模行为动态,提前预警潜在危险,提升安全性和任务完成率。

Comments Accepted by the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05842 2026-08-04 cs.RO 版本更新 57%

Expert Knowledge-driven Reinforcement Learning for Autonomous Racing via Trajectory Guidance and Dynamics Constraints

专家知识驱动的强化学习用于自动驾驶赛车:通过轨迹引导和动力学约束

Bo Leng, Weiqi Zhang, Zhuoren Li, Lu Xiong, Guizhe Jin, Ran Yu, Chen Lv

机构 * College of Automotive and Energy Engineering, Tongji University(同济大学汽车与能源工程学院) School of Mechanical and Aerospace Engineering, Nanyang Technological University(南洋理工大学机械与航空航天工程学院)

专题命中 规划控制 :autonomous driving(abstract);分类 cs.RO

AI总结 本文提出TraD-RL方法,通过轨迹引导和动力学约束强化学习,提升自动驾驶赛车的性能与安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.00853 2026-08-04 cs.SE cs.SY eess.SY 版本更新 50%

Guidance on the Safety Assurance of Autonomous Systems in Complex Environments (SACE)

复杂环境下自主系统安全保障指南(SACE)

Richard Hawkins, Rob Alexander, Matt Osborne, Mike Parsons, Mark Nicholson, John McDermid, Ibrahim Habli

专题命中 规划控制 :autonomous driving(abstract)

AI总结 该研究提出SACE方法,通过安全论证模式与流程,将安全保障整合到自主系统开发中,生成安全可接受性证据,解决复杂环境下自主系统的安全保障问题。

Comments Version 1.1

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 端到端驾驶 1 篇

2602.10719 2026-08-04 cs.RO cs.CV 版本更新 76%

From Representational Complementarity to Dual Systems: Synergizing VLM and Vision-Only Backbones for End-to-End Driving

从表征互补性到双系统:协同VLM和纯视觉骨干网络用于端到端驾驶

Sining Ang, Yuguang Yang, Chenxu Dang, Canyu Chen, Cheng Chi, Haiyan Liu, Xuanyao Mao, Jason Bao, Xuliang, Bingchuan Sun, Yan Wang

机构 * Department of Automation, University of Science and Technology of China(中国科学技术大学自动化系) School of Electronic Information Engineering, Beihang University(北京航空航天大学电子信息工程学院) School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院) National Superior College for Engineers, Beihang University(北京航空航天大学国家级工程师学院) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Lenovo Group Limited(联想集团有限公司) Institute for AI Industry Research, Tsinghua University(清华大学人工智能产业研究院)

专题命中 端到端驾驶 :end-to-end driving(title);分类 cs.RO、cs.CV

AI总结 本文通过协同VLM和纯视觉骨干网络,提出HybridDriveVLA和DualDriveVLA,提升端到端驾驶的PDMS性能。

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 激光雷达 2 篇

2505.18819 2026-08-04 cs.CV 版本更新 70%

Parameter-Efficient CLIP Adaptation for 3D Understanding via Unified Tokenization

用于3D理解的参数高效CLIP适配:通过统一分词实现

Guofeng Mei, Qinfeng Xiao, Bin Ren, Luigi Riz, Juan Liu, Xiaoshui Huang, Xu Zheng, Nicu Sebe, Ming-Hsuan Yang, Fabio Poiesi

机构 * Fondazione Bruno Kessler(布鲁诺·科塞拉基金会) University of Trento(特伦托大学) University of Pisa(比萨大学) Beijing Forestry University(北京林业大学) Shanghai Jiao Tong University(上海交通大学) Hong Kong University of Science and Technology (GZ)(香港科技大学) Shandong University(山东大学) University of California, Merced(加州大学默塞德分校)

专题命中 激光雷达 :LiDAR(abstract,abstract_cn);分类 cs.CV

AI总结 本文提出参数高效框架UTok3D,通过学习尺度归一化3D分词器,实现冻结CLIP视觉主干在无标注情况下对不同尺度点云的复用,完成3D分割任务。

Comments 14 pages, tokenizer

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08071 2026-08-04 cs.IT math.IT 版本更新 67%

Integrated Localization, Mapping, and Communication through VCSEL-Based Light-emitting RIS (LeRIS)

基于垂直腔面发射激光器(VCSEL)的发光可重构智能表面(LeRIS)实现定位、建图与通信的集成

Rashid Iqbal, Dimitrios Bozanis, Dimitrios Tyrovolas, Sotiris Ioannidis, Christos K. Liaskos, Muhammad Ali Imran, George K. Karagiannidis, Hanaa Abumarshoud

专题命中 激光雷达 :LiDAR(abstract,abstract_cn)

AI总结 本文提出基于VCSEL的LeRIS框架,可联合实现用户定位、避障建图与毫米波通信,通过仿真验证其能达到厘米级定位精度、可靠障碍物检测及频谱效率与用户速率增益。

详情

展开后加载摘要…

URL PDF HTML 收藏