arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

2026-04-29 至 2026-04-29 共收录 9 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 6 篇

2604.24622 2026-04-29 cs.CV cs.AI 93%

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies

CF-VLA:面向视觉-语言-动作策略的高效粗到细动作生成

Fan Du, Feng Yan, Jianxiong Wu, Xinrun Xu, Weiye Zhang, Weinong Wang, Yu Guo, Bin Qian, Zhihai He, Fei Wang, Heng Yang

机构 * Southern University of Science and Technology(南方科技大学) Xi’an Jiaotong University(西安交通大学) United Nova Technology(联合Nova技术) University of Science and Technology of China(中国科学技术大学)

专题命中 VLA模型 :VLA(title,title_cn);vision-language-action(title,abstract);action model(abstract);分类 cs.CV、cs.AI

AI总结 本文提出CF-VLA,通过粗到细两阶段方法优化视觉-语言-动作策略的动作生成,提升效率与性能,在低NFE条件下实现更优的效率-性能平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29844 2026-04-29 cs.RO cs.AI cs.CV cs.LG 92%

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA

DIAL: 通过潜在世界建模解耦意图与动作以实现端到端VLA

Yi Chen, Yuying Ge, Hui Zhou, Mingyu Ding, Yixiao Ge, Xihui Liu

机构 * The University of Hong Kong(香港大学) XPENG Robotics(小鹏机器人) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 VLA模型 :VLA(title,title_cn);vision-language-action(abstract);分类 cs.RO、cs.CV、cs.AI

AI总结 DIAL通过潜在意图瓶颈解耦意图与动作,利用VLM进行潜在世界建模并结合轻量策略实现端到端VLA,实验表明其在RoboCasa GR1任务中优于现有方法,且在真实世界部署中表现稳健。

Comments Project page: https://xpeng-robotics.github.io/dial

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24921 2026-04-29 cs.RO cs.AI cs.CL cs.CV 91%

Libra-VLA: Achieving Learning Equilibrium via Asynchronous Coarse-to-Fine Dual-System

Libra-VLA:通过异步粗到细双系统实现学习平衡

Yifei Wei, Linqing Zhong, Yi Liu, Yuxiang Lu, Xindong He, Maoqing Yao, Guanghui Ren

机构 * Beihang University(北航) AgiBot

专题命中 VLA模型 :VLA(title,title_cn);vision-language-action(abstract);分类 cs.RO、cs.CV、cs.AI

AI总结 Libra-VLA通过粗到细双系统架构,解决视觉-语言-动作模型在机器人操作中语义与动作间的鸿沟问题,实现训练平衡与高效执行。

Comments Accepted to the Main Conference of ACL 2026. Project page: https://libra-vla.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11075 2026-04-29 cs.RO 77%

RISE: Self-Improving Robot Policy with Compositional World Model

RISE:具有组合世界模型的自改进机器人策略

Jiazhi Yang, Kunyang Lin, Jinwei Li, Wencong Zhang, Tianwei Lin, Longyan Wu, Zhizhong Su, Hao Zhao, Ya-Qin Zhang, Li Chen, Ping Luo, Xiangyu Yue, Hongyang Li

机构 * The Chinese University of Hong Kong(香港中文大学) Kinetix AI The University of Hong Kong(香港大学) Shanghai Innovation Institute(上海创新研究院) Horizon Robotics Tsinghua University(清华大学)

专题命中 VLA模型 :VLA(abstract,abstract_cn);vision-language-action(abstract);分类 cs.RO

AI总结 本文提出RISE框架,通过组合世界模型提升机器人策略鲁棒性,在动态砖块排序、背包打包和箱盖闭合任务中分别取得+35%、+45%和+35%的性能提升。

Comments RSS 2026. Project page: https://opendrivelab.com/RISE/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22347 2026-04-29 astro-ph.GA 67%

A Morphological Identification and Study of Radio Galaxies from LoTSS DR2. I. The "Winged'' Radio Galaxies

从LoTSS DR2对无线电星系的形态识别与研究。I. '有翼'的无线电星系

Soumen Kumar Bera, Taotao Fang, Tapan K. Sasmal, M. Kunert-Bajraszewska, Xuelei Chen, Soumen Mondal

专题命中 VLA模型 :VLA(abstract,abstract_cn)

AI总结 利用LoTSS DR2最新高分辨率数据,研究了无线电星系的形态分类,发现621个新'有翼'星系及403个候选体,其中382个为X型,239个为Z型,大部分为FR-II类,具有高平均无线电功率和边缘增强的射电晕。

Journal ref 2025ApJS..278...34B

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24796 2026-04-29 q-bio.OT cs.LG 57%

A multi-stage soft computing framework for complex disease modelling and decision support: A liver cirrhosis case study

一种多阶段软计算框架用于复杂疾病建模和决策支持:肝硬化案例研究

Xueyuan Huang, Yuheng Wang, Yuanzhi He, Siqi Gou, Lu Bai, Wenqian Wu, Peifeng Liu, Aijia Wang, Tianhui Fan, Ze Zhou, Jiayu Xu

机构 * organization= Department of Hepatobiliary Surgery, the Second Affiliated Hospital of Chongqing Medical University , city= Chongqing , postcode= 400010 , country= China organization= State Key Laboratory of Respiratory Health Multimorbidity, Institute of Basic Medical Sciences \& School of Basic Medicine, Chinese Academy of Medical Sciences \& Peking Union Medical College , city= Beijing , postcode= 100005 , country= China organization= School of Computer Science Informatics, Cardiff University , city= Wales , postcode= CF24 0DE , country= United Kingdom organization= Department of Medical Oncology, the First Hospital of China Medical University , city= Shenyang , postcode= 110001 , country= China organization= Institute for Innovation Development, Tsinghua University , city= Beijing , postcode= 100086 , country= China organization= Department of Liver Surgery, Peking Union Medical College Hospital, Chinese Academy of Medical Sciences \& Peking Union Medical College , city= Beijing , country= China

专题命中 VLA模型 :action model(abstract);分类 cs.LG

AI总结 本文提出一种多阶段软计算框架,用于复杂疾病建模与治疗探索,通过整合单细胞转录组数据、高维网络特征稳定化、多模型学习和深度表征学习,提升肝硬化等复杂疾病的诊断与治疗决策能力。

Comments 20 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 机器人基础模型 1 篇

2511.16518 2026-04-29 cs.RO cs.CL cs.CV 81%

MiMo-Embodied: X-Embodied Foundation Model Technical Report

MiMo-Embodied:X-Embodied基础模型技术报告

Xiaoshuai Hao, Lei Zhou, Zhijian Huang, Zhiwen Hou, Yingbo Tang, Lingfeng Zhang, Guang Li, Zheng Lu, Shuhuai Ren, Xianhui Meng, Yuchen Zhang, Jing Wu, Jinghui Lu, Chenxu Dang, Jiayi Guan, Jianhua Wu, Zhiyi Hou, Hanbing Li, Shumeng Xia, Mingliang Zhou, Yinan Zheng, Zihao Yue, Shuhao Gu, Hao Tian, Yuannan Shen, Jianwei Cui, Wen Zhang, Shaoqing Xu, Bing Wang, Haiyang Sun, Zeyu Zhu, Yuncheng Jiang, Zibin Guo, Chuhong Gong, Chaofan Zhang, Wenbo Ding, Kun Ma, Guang Chen, Rui Cai, Diyun Xiang, Heng Qu, Fuli Luo, Hangjun Ye, Long Chen

机构 * Xiaomi Embodied Intelligence Team(小米具身智能团队)

专题命中 机器人基础模型 :embodied foundation model(title,abstract);分类 cs.RO、cs.CV

AI总结 MiMo-Embodied是首个跨具身基础模型,成功整合并实现自动驾驶和具身AI领域的最先进性能,在17个具身AI基准和12个自动驾驶基准中均取得新纪录,通过多阶段学习、数据构建和CoT/RL微调实现领域间的强正迁移。

Comments Code: https://github.com/XiaomiMiMo/MiMo-Embodied | Model: https://huggingface.co/XiaomiMiMo/MiMo-Embodied-7B

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 数据集与评测 1 篇

2601.03136 2026-04-29 cs.CL cs.AI cs.RO 79%

Limited Linguistic Diversity in Embodied AI Datasets

具身AI数据集中的语言多样性有限

Selma Wanna, Agnes Luhtaru, Jonathan Salfity, Ryan Barron, Juston Moore, Cynthia Matuszek, Mitch Pryor

机构 * Los Alamos National Laboratory(洛斯阿拉莫斯国家实验室) Institute of Computer Science, University of Tartu(塔尔图大学计算机科学学院) Department of Mechanical Engineering, The University of Texas at Austin(德克萨斯大学奥斯汀分校机械工程系) University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校)

专题命中 数据集与评测 :VLA(abstract,abstract_cn);vision-language-action(abstract);分类 cs.RO、cs.AI

AI总结 本文分析了多个广泛使用的视觉-语言-动作数据集的语言特性,发现其指令存在高度重复和结构单一的问题,旨在推动更详细的 dataset 报告和语言覆盖的扩展策略。

Comments Accepted to ACL 2026 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 部署与泛化 1 篇

2507.14245 2026-04-29 cs.LG cond-mat.mtrl-sci cs.AI cs.CE q-bio.BM 62%

Curriculum-guided multimodal representation learning enables generalizable prediction of nanomaterial-protein interactions

基于课程引导的多模态表示学习可实现纳米材料-蛋白质相互作用的通用预测

Hengjie Yu, Kenneth A. Dawson, Haiyun Yang, Shuya Liu, Yan Yan, Yaochu Jin

机构 * School of Engineering, Westlake University(西湖大学工程学院) Institute of Advanced Technology, Westlake Institute for Advanced Study(西湖研究所在先进科技研究院) Centre for BioNano Interactions, School of Chemistry, University College Dublin(都柏林大学学院化学系生物纳米相互作用中心) School of Biomolecular and Biomedical Science, UCD Conway Institute of Biomolecular and Biomedical Research(都柏林大学学院生物分子与生物医学科学系,康沃利斯生物分子与生物医学研究学院)

专题命中 部署与泛化 :action model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出CuMMI模型,通过多阶段课程学习和大规模数据集,实现纳米材料-蛋白质相互作用的通用预测,验证了其在不同数据集上的鲁棒性和可迁移性。

Comments 36 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏