arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

共收录 9819 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 9146 篇

2601.04223 2026-01-09 cs.CY cs.AI cs.LG econ.GN q-fin.EC stat.ME 62%

Beyond Interaction Effects: Two Logics for Studying Population Inequalities

超越交互效应:研究人口不平等的两种逻辑

Adel Daoud

机构 * Institute for Analytical Sociology, Linköping University, Sweden(分析社会学研究所,利厄普斯大学,瑞典) Division of Data Science and AI, Chalmers University of Technology, Sweden(数据科学与人工智能部门,查尔姆斯理工大学,瑞典) The AI and Development Lab, www.aidevlab.org(人工智能与发展实验室)

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出一种框架,用于在演绎逻辑和归纳逻辑之间导航,探讨研究人口不平等时可解释性与灵活性的权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23037 2026-01-01 cs.AI cs.CL cs.LG 62%

Agentic Large Language Models, a survey

代理大语言模型:综述

Aske Plaat, Max van Duijn, Niki van Stein, Mike Preuss, Peter van der Putten, Kees Joost Batenburg

机构 * Leiden University Leiden Netherlands Leiden University \& AI Lab, Pegasystems Leiden Netherlands Leiden University Leiden University \& AI Lab, Pegasystems

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

AI总结 本文综述了代理大语言模型的研究现状,探讨了其在医疗诊断、物流和金融分析等领域的应用,并提出通过推理、行动和交互提升大语言模型能力的未来研究方向。

Comments Website: https://askeplaat.github.io/agentic-llm-survey-site/

Journal ref JAIR volume 84, article 29, December 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22466 2025-12-30 cs.LG cs.AI 62%

AMBIT: Augmenting Mobility Baselines with Interpretable Trees

AMBIT:通过可解释的树模型增强移动性基线

Qizhi Wang

机构 * PingCAP, Data & AI-Innovation Lab(PingCAP数据与人工智能创新实验室)

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

AI总结 AMBIT通过结合可解释的树模型和物理基线,提升OD流量预测的准确性和可解释性,为城市决策提供支持。

Comments 15 pages; 12 figures; 30 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20052 2025-12-24 cs.AI cs.RO 62%

Learning Skills from Action-Free Videos

从无动作视频中学习技能

Hung-Chieh Fang, Kuo-Han Hung, Chu-Rong Chen, Po-Jung Chou, Chun-Kai Yang, Po-Chen Ko, Yu-Chiang Wang, Yueh-Hua Wu, Min-Hung Chen, Shao-Hua Sun

专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.AI

AI总结 本文提出SOF框架,通过从无动作视频中学习潜在技能,实现视频动态与机器人动作的对齐,提升多任务和长周期任务的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19024 2025-12-23 cs.RO cs.AI 62%

IndoorUAV: Benchmarking Vision-Language UAV Navigation in Continuous Indoor Environments

IndoorUAV:在连续室内环境中进行视觉-语言无人机导航的基准测试

Xu Liu, Yu Liu, Hanshuo Qiu, Yang Qirong, Zhouhui Lian

专题命中 VLA模型 :VLA(abstract);分类 cs.RO、cs.AI

AI总结 IndoorUAV通过构建室内无人机视觉-语言导航基准,结合多模态推理和任务分解,为长视距和短视距导航提供高质量数据与模型支持。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05066 2025-12-05 cs.LG cs.AI cs.CL 62%

Multi-LLM Collaboration for Medication Recommendation

多LLM协作用于药物推荐

Huascar Sanchez, Briland Hitaj, Jules Bergmann, Linda Briesemeister

机构 * Computer Science Laboratory, SRI International(SRI国际计算机科学实验室) University of Maryland St. Joseph Medical Center(马里兰大学圣约瑟夫医疗中心)

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出基于LLM化学的多模型协作方法,通过增强互补性、稳定性和校准性,提高药物推荐的可靠性与可信度。

Comments 8 pages, 5 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01629 2025-12-03 cs.CV cs.RO 62%

SPARK: Sim-ready Part-level Articulated Reconstruction with VLM Knowledge

SPARK: 面向模拟的关节化重建与VLK知识

Yumeng He, Ying Jiang, Jiayin Lu, Yin Yang, Chenfanfu Jiang

专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.CV

AI总结 SPARK通过结合VLK和生成扩散模型,实现从单张图像中生成物理一致的关节化3D物体,提升机器人操作和交互建模的应用效果。

Comments Project page: https://heyumeng.com/SPARK/index.html. 17 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11726 2025-12-02 cs.LG cs.AI 62%

Saga: Capturing Multi-granularity Semantics from Massive Unlabelled IMU Data for User Perception

Saga:从大量未标记IMU数据中捕捉多粒度语义以实现用户感知

Yunzhe Li, Facheng Hu, Hongzi Zhu, Shifan Zhang, Liang Zhang, Shan Chang, Minyi Guo

机构 * Shanghai Jiao Tong University(上海交通大学) Donghua University(东华大学)

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

AI总结 Saga通过预训练和贝叶斯优化实现从大量未标记IMU数据中捕捉多粒度语义,仅用少量标记数据即可达到高用户感知准确率。

Comments 2025 IEEE 45th International Conference on Distributed Computing Systems (ICDCS)

Journal ref Proceedings of the IEEE International Conference on Distributed Computing Systems (ICDCS), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02549 2025-12-01 cs.CV cs.RO 62%

MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming

MonoDream:基于全景梦境的单目视觉-语言导航

Shuo Wang, Yongcai Wang, Zhaoxin Fan, Yucheng Wang, Maiyue Chen, Kaihui Wang, Zhizhong Su, Wanting Li, Xudong Cai, Yeying Jin, Deying Li

机构 * Horizon Robotics

专题命中 VLA模型 :VLA(abstract);分类 cs.RO、cs.CV

AI总结 MonoDream通过轻量级VLA框架和潜在全景梦境任务,提升单目视觉-语言导航的性能,缩小与全景方法的差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03806 2025-11-18 cs.RO cs.LG 62%

Certified Coil Geometry Learning for Short-Range Magnetic Actuation and Spacecraft Docking Application

Yuta Takahashi, Hayate Tajima, Shin-ichiro Sakai

机构 * Department of Mechanical Engineering, Institute of Science Tokyo(科学东京研究院机械工程系) Satellite Research and Development, Interstellar Technologies Inc.(星际技术公司卫星研发部) Department of Advanced Energy, The University of Tokyo(东京大学先进能源系) Department of Spacecraft Engineering, Japan Aerospace Exploration Agency(日本宇宙航空研究开发机构航天器工程系)

专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.LG

Comments Submitted to IEEE Robotics and Automation Letters

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05210 2025-11-10 cs.LG cs.AI cs.SY eess.SY 62%

Advanced Hybrid Transformer LSTM Technique with Attention and TS Mixer for Drilling Rate of Penetration Prediction

Saddam Hussain Khan

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

Comments 31 Pages, 16 Figures, 9 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23189 2025-10-28 cs.CL cs.AI cs.LG 62%

DREaM: Drug-Drug Relation Extraction via Transfer Learning Method

Ali Fata, Hossein Rahmani, Parinaz Soltanzadeh, Amirhossein Derakhshan, Behrouz Minaei Bidgoli

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13921 2025-10-17 cs.RO cs.AI 62%

APEX: Empowering LLMs with Physics-Based Task Planning for Real-time Insight

Wanjing Huang, Weixiang Yan, Zhen Zhang, Ambuj Singh

专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16815 2025-10-15 cs.CV cs.RO 62%

Image Quality Assessment for Embodied AI

Chunyi Li, Jiaohao Xiao, Jianbo Zhang, Farong Wen, Zicheng Zhang, Yuan Tian, Xiangyang Zhu, Xiaohong Liu, Zhengxue Cheng, Weisi Lin, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Lab(上海人工智能实验室) Nanyang Technological University(南洋理工大学)

专题命中 VLA模型 :vision language action(abstract);分类 cs.RO、cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07313 2025-10-09 cs.CV cs.RO 62%

WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation

Zezhong Qian, Xiaowei Chi, Yuming Li, Shizun Wang, Zhiyuan Qin, Xiaozhu Ju, Sirui Han, Shanghang Zhang

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(信息处理国家重点实验室,计算机学院,北京大学) Hong Kong University of Science and Technology(香港科技大学) National University of Singapore(新加坡国家大学) Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心)

专题命中 VLA模型 :VLA(abstract);分类 cs.RO、cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19227 2025-09-24 cs.CV cs.AI 62%

MsFIN: Multi-scale Feature Interaction Network for Traffic Accident Anticipation

Tongshuai Wu, Chao Lu, Ze Song, Yunlong Lin, Sizhe Fan, Xuemei Chen

机构 * School of Mechanical Engineering, Beijing Institute of Technology(机械工程学院,北京理工大学)

专题命中 VLA模型 :action model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17888 2025-09-23 cs.CV cs.AI 62%

Trainee Action Recognition through Interaction Analysis in CCATT Mixed-Reality Training

Divya Mereddy, Marcos Quinones-Grueiro, Ashwin T S, Eduardo Davalos, Gautam Biswas, Kent Etherton, Tyler Davis, Katelyn Kay, Jill Lear, Benjamin Goldberg

专题命中 VLA模型 :action model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15968 2025-09-22 cs.RO cs.CV 62%

CoReVLA: A Dual-Stage End-to-End Autonomous Driving Framework for Long-Tail Scenarios via Collect-and-Refine

Shiyu Fang, Yiming Cui, Haoyang Liang, Chen Lv, Peng Hang, Jian Sun

机构 * College of Transportation, Tongji University(同济大学交通运输学院)

专题命中 VLA模型 :VLA(abstract);分类 cs.RO、cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15882 2025-09-22 cs.CV cs.AI 62%

Self-Supervised Cross-Modal Learning for Image-to-Point Cloud Registration

Xingmei Wang, Xiaoyu Hu, Chengkai Huang, Ziyan Zeng, Guohao Nie, Quan Z. Sheng, Lina Yao

专题命中 VLA模型 :action model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04076 2025-09-17 cs.RO cs.AI 62%

Keypoint-based Diffusion for Robotic Motion Planning on the NICOL Robot

Lennart Clasmeier, Jan-Gerrit Habekost, Connor Gäde, Philipp Allgeuer, Stefan Wermter

机构 * Dept. of Informatics, University of Hamburg(信息学院,汉堡大学)

专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.AI

Comments Accepted and published at the 34th International Conference on Artificial Neural Networks (ICANN 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15214 2025-09-16 cs.RO cs.LG 62%

Think Small, Plan Smart: Minimalist Symbolic Abstraction and Heuristic Subspace Search for LLM-Guided Task Planning

Junfeng Tang, Yuping Yan, Zihan Ye, Zhenshou, Song, Zeqi Zheng, Yaochu Jin

机构 * Zhejiang University(浙江大学) Westlake University(西湖大学) UCAS-Terminus AI Lab(UCAS-terminus人工智能实验室) Victoria University of Wellington(威灵顿维多利亚大学)

专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14799 2025-08-28 cs.HC cs.AI cs.LG 62%

Analyzing Character Representation in Media Content using Multimodal Foundation Model: Effectiveness and Trust

Evdoxia Taka, Debadyuti Bhattacharya, Joanne Garde-Hansen, Sanjay Sharma, Tanaya Guha

机构 * University of Glasgow(格拉斯哥大学) University of Leeds(利兹大学) University of Warwick(沃里克大学)

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06538 2025-07-30 cs.RO cs.AI 62%

OPAL: Encoding Causal Understanding of Physical Systems for Robot Learning

Daniel Tcheurekdjian, Joshua Klasmeier, Tom Cooney, Christopher McCann, Tyler Fenstermaker

专题命中 VLA模型 :vision-language-action(abstract);分类 cs.RO、cs.AI

Comments We withdraw our submission following peer review feedback that identified methodological limitations: specifically, our experimental design does not adequately support the causal claims made in the submission. The work was preliminary undergraduate research that requires substantial additional experimental validation to properly establish the proposed causal relationships

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16824 2025-07-22 cs.CV cs.AI 62%

PerspectiveNet: Multi-View Perception for Dynamic Scene Understanding

Vinh Nguyen

机构 * Uppsala University(乌普萨拉大学) Florida Institute of Technology(佛罗里达理工学院)

专题命中 VLA模型 :action model(abstract);分类 cs.CV、cs.AI

Comments 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15468 2025-06-19 cs.HC cs.AI cs.LG 62%

Co-Creative Learning via Metropolis-Hastings Interaction between Humans and AI

Ryota Okumura, Tadahiro Taniguchi, Akira Taniguchi, Yoshinobu Hagiwara

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09940 2025-06-12 cs.LG cs.AI stat.ML 62%

The Sample Complexity of Online Strategic Decision Making with Information Asymmetry and Knowledge Transportability

Jiachen Hu, Rui Ai, Han Zhong, Xiaoyu Chen, Liwei Wang, Zhaoran Wang, Zhuoran Yang

机构 * School of Computer Science, Peking University(北京大学计算机科学学院) Center for Data Science, Peking University(北京大学数据科学中心) National Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University(北京大学通用人工智能国家重点实验室) Northwestern University(西北大学) Yale University(耶鲁大学)

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

Comments Accepted at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07286 2025-06-06 q-bio.BM cs.AI cs.LG 62%

Piloting Structure-Based Drug Design via Modality-Specific Optimal Schedule

Keyue Qiu, Yuxuan Song, Zhehuan Fan, Peidong Liu, Zhe Zhang, Mingyue Zheng, Hao Zhou, Wei-Ying Ma

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

Comments Accepted to ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23791 2025-06-02 cs.CR cs.AI cs.LG 62%

Evaluating Query Efficiency and Accuracy of Transfer Learning-based Model Extraction Attack in Federated Learning

Sayyed Farid Ahamed, Sandip Roy, Soumya Banerjee, Marc Vucovich, Kevin Choi, Abdul Rahman, Alison Hu, Edward Bowen, Sachin Shetty

机构 * Center for Secure & Intelligent Critical Systems, Old Dominion University, Virginia, USA(安全与智能关键系统中心,旧 Dominion 大学,弗吉尼亚州,美国) School of Cybersecurity, Old Dominion University, Virginia, USA(网络安全学院,旧 Dominion 大学,弗吉尼亚州,美国)

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

Comments Accepted at IEEE IWCMC. 6 pages, 4 Figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01932 2025-05-29 cs.RO cs.LG 62%

Bridging Language, Vision and Action: Multimodal VAEs in Robotic Manipulation Tasks

Gabriela Sejnova, Michal Vavrecka, Karla Stepanova

机构 * Czech Institute of Informatics, Robotics and Cybernetics(捷克信息学、机器人学与自动控制研究所) Czech Technical University in Prague(布拉格捷克技术大学)

专题命中 VLA模型 :vision-language-action(abstract);分类 cs.RO、cs.LG

Comments 7 pages, 5 figures, 2 tables, conference

Journal ref 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18010 2025-05-28 cs.RO cs.AI cs.HC 62%

Sky-Drive: A Distributed Multi-Agent Simulation Platform for Human-AI Collaborative and Socially-Aware Future Transportation

Zilin Huang, Zihao Sheng, Zhengyang Wan, Yansong Qu, Yuhao Luo, Boyue Wang, Pei Li, Yen-Jung Chen, Jiancong Chen, Keke Long, Jiayi Meng, Yue Leng, Sikai Chen

机构 * Department of Civil and Environmental Engineering, University of Wisconsin-Madison(土木与环境工程系,威斯康星大学麦迪逊分校) Lyles School of Civil and Construction Engineering, Purdue University(莱尔斯土木与建设工程学院,普渡大学) Elmore Family School of Electrical and Computer Engineering, Purdue University(埃尔摩家族电气与计算机工程学院,普渡大学) Department of Computer Science and Engineering, The University of Texas at Arlington(计算机科学与工程系,德克萨斯大学阿灵顿分校) Google(谷歌)

专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.AI

Comments 14 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏