arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

共收录 9819 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 9146 篇

2509.21607 2025-09-29 cs.LG 57%

Causal Abstraction Inference under Lossy Representations

Kevin Xia, Elias Bareinboim

机构 * CausalAI Lab, Columbia University(因果推理实验室,哥伦比亚大学)

专题命中 VLA模型 :action model(abstract);分类 cs.LG

Comments 35 pages, 8 figures, published at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19719 2025-09-25 cs.CV 57%

Frequency-domain Multi-modal Fusion for Language-guided Medical Image Segmentation

Bo Yu, Jianhua Yang, Zetao Du, Yan Huang, Chenglong Li, Liang Wang

机构 * School of Computer Science and Technology, Anhui University(安徽大学计算机科学与技术学院) NLPR, MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Information Science and Technology, ShanghaiTech University(上海科技大学信息科学与技术学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) School of Artificial Intelligence, Anhui University(安徽大学人工智能学院)

专题命中 VLA模型 :action model(abstract);分类 cs.CV

Comments Accepted by MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07358 2025-09-22 cs.RO 57%

SymBridge: A Human-in-the-Loop Cyber-Physical Interactive System for Adaptive Human-Robot Symbiosis

Haoran Chen, Yiteng Xu, Yiming Ren, Yaoqin Ye, Xinran Li, Ning Ding, Yuxuan Wu, Yaoze Liu, Peishan Cong, Ziyi Wang, Bushi Liu, Yuhan Chen, Zhiyang Dou, Xiaokun Leng, Manyi Li, Yuexin Ma, Changhe Tu

机构 * Shandong University(山东大学) ShanghaiTech University(上海科技大学) The University of Hong Kong(香港大学) LEJU(Shenzhen) Robotics Co., Ltd(LEJU(深圳)机器人有限公司)

专题命中 VLA模型 :action model(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02445 2025-09-08 cs.CV 57%

Towards High-Fidelity, Identity-Preserving Real-Time Makeup Transfer: Decoupling Style Generation

Lydia Kin Ching Chau, Zhi Yu, Ruowei Jiang

专题命中 VLA模型 :action model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03939 2025-09-05 cs.CR cs.LG 57%

LMAE4Eth: Generalizable and Robust Ethereum Fraud Detection by Exploring Transaction Semantics and Masked Graph Embedding

Yifan Jia, Yanbin Wang, Jianguo Sun, Ye Tian, Peng Qian

机构 * Yantai Research Institute, Harbin Engineering University(哈尔滨工程大学烟台研究院) Hangzhou Research Institute, Xidian University(西安电子科技大学杭州研究院) College of Computer Science, Zhejiang University(浙江大学计算机学院)

专题命中 VLA模型 :action model(abstract);分类 cs.LG

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03890 2025-09-05 cs.AI 57%

FaMA: LLM-Empowered Agentic Assistant for Consumer-to-Consumer Marketplace

Yineng Yan, Xidong Wang, Jin Seng Cheng, Ran Hu, Wentao Guan, Nahid Farahmand, Hengte Lin, Yue Li

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校) Meta Platforms, Inc(Meta平台公司)

专题命中 VLA模型 :action model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07897 2025-09-04 cs.LG cs.IR cs.MA cs.SY eess.SY 57%

The Nah Bandit: Modeling User Non-compliance in Recommendation Systems

Tianyue Zhou, Jung-Hoon Cho, Cathy Wu

机构 * Department of Civil and Environmental Engineering(土木与环境工程系) Laboratory for Information & Decision Systems(信息与决策系统实验室) Massachusetts Institute of Technology(麻省理工学院) Institute for Data, Systems, and Society(数据、系统与社会研究所)

专题命中 VLA模型 :action model(abstract);分类 cs.LG

Comments 12 pages, 8 figures, accepted by IEEE Transactions on Control of Network Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09625 2025-08-27 cs.CV cs.CL 57%

Weakly-Supervised 3D Visual Grounding based on Visual Language Alignment

Xiaoxu Xu, Yitian Yuan, Qiudan Zhang, Wenhui Wu, Zequn Jie, Lin Ma, Xu Wang

机构 * College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) Meituan Inc.(美团公司) College of Electronics and Information Engineering, Shenzhen University(深圳大学电子与信息工程学院) Guangdong Key Laboratory of Intelligent Information Processing(广东省智能信息处理重点实验室)

专题命中 VLA模型 :VLA(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10889 2025-08-22 cs.AI 57%

On Learning Action Costs from Input Plans

Marianela Morales, Alberto Pozanco, Giuseppe Canonaco, Sriram Gopalakrishnan, Daniel Borrajo, Manuela Veloso

机构 * J.P. Morgan AI Research(摩根大通人工智能研究)

专题命中 VLA模型 :action model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.10099 2025-08-20 cs.CV 57%

STGFormer: Spatio-Temporal GraphFormer for 3D Human Pose Estimation in Video

Yang Liu, Zhiyong Zhang

机构 * Sun Yat-sen University, School of Electronics and Communication Engineering(中山大学电子与通信工程学院)

专题命中 VLA模型 :action model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16752 2025-08-19 cs.IR cs.AI 57%

Action is All You Need: Dual-Flow Generative Ranking Network for Recommendation

Hao Guo, Erpeng Xue, Lei Huang, Shichao Wang, Xiaolei Wang, Lei Wang, Jinpeng Wang, Sheng Chen

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院) Tsinghua University(清华大学)

专题命中 VLA模型 :action model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11517 2025-08-18 cs.CV 57%

A Real-time Concrete Crack Detection and Segmentation Model Based on YOLOv11

Shaoze Huang, Qi Liu, Chao Chen, Yuhang Chen

机构 * MSE, AHUT, China(机械与电子工程学院,哈尔滨工业大学(哈尔滨)) IoT, AHUT, China(物联网学院,哈尔滨工业大学(哈尔滨)) IST, SDUST, China(信息科学与技术学院,山东大学(济南))

专题命中 VLA模型 :action model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03936 2025-08-14 cs.CV 57%

Learning Adaptive Node Selection with External Attention for Human Interaction Recognition

Chen Pang, Xuequan Lu, Qianyu Zhou, Lei Lyu

机构 * Shandong Normal University(山东师范大学) University of Western Australia(西澳大学) Jilin University(吉林大学) University of Western(西澳大学)

专题命中 VLA模型 :action model(abstract);分类 cs.CV

Comments Accepted by ACM MM25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23778 2025-08-13 cs.CV 57%

Half-Physics: Enabling Kinematic 3D Human Model with Physical Interactions

Li Siyao, Yao Feng, Omid Taheri, Chen Change Loy, Michael J. Black

机构 * Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所) S-Lab, Nanyang Technological University(南洋理工大学S实验室) Stanford University(斯坦福大学)

专题命中 VLA模型 :action model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07318 2025-08-12 cs.CV 57%

RORPCap: Retrieval-based Objects and Relations Prompt for Image Captioning

Jinjing Gu, Tianbao Qin, Yuanyuan Pu, Zhengpeng Zhao

机构 * School of Information Science and Engineering(信息科学与工程学院)

专题命中 VLA模型 :action model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08333 2025-08-12 cs.CV 57%

DAViD: Modeling Dynamic Affordance of 3D Objects Using Pre-trained Video Diffusion Models

Hyeonwoo Kim, Sangwon Baik, Hanbyul Joo

机构 * Seoul National University(首尔国立大学)

专题命中 VLA模型 :action model(abstract);分类 cs.CV

Comments Project Page: https://snuvclab.github.io/david/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05799 2025-08-11 cs.SE cs.AI cs.HC 57%

AI-Guided Exploration of Large-Scale Codebases

Yoseph Berhanu Alebachew

机构 * Department of Computer Science(计算机科学系) Virginia Tech(弗吉尼亚理工大学)

专题命中 VLA模型 :action model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00469 2025-08-08 cs.SD cs.AI cs.IR eess.AS 57%

MIRFLEX: Music Information Retrieval Feature Library for Extraction

Anuradha Chopra, Abhinaba Roy, Dorien Herremans

专题命中 VLA模型 :action model(abstract);分类 cs.AI

Comments 2 pages, 4 tables, submitted to Extended Abstracts for the Late-Breaking Demo Session of the 25th Int. Society for Music Information Retrieval Conf., San Francisco, United States, 2024

Journal ref ISMIR Demo paper, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21045 2025-08-05 cs.CV 57%

Reconstructing 4D Spatial Intelligence: A Survey

Yukang Cao, Jiahao Lu, Zhisheng Huang, Zhuowen Shen, Chengfeng Zhao, Fangzhou Hong, Zhaoxi Chen, Xin Li, Wenping Wang, Yuan Liu, Ziwei Liu

机构 * S-Lab, College of Computing and Data Science, Nanyang Technological University(S实验室,计算与数据科学学院,南洋理工大学) Intelligent Graphics Lab, The Hong Kong University of Science and Technology(智能图形实验室,香港科学与技术大学)

专题命中 VLA模型 :action model(abstract);分类 cs.CV

Comments Project page: https://github.com/yukangcao/Awesome-4D-Spatial-Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00580 2025-07-29 cs.RO 57%

MRHaD: Mixed Reality-based Hand-Drawn Map Editing Interface for Mobile Robot Navigation

Takumi Taki, Masato Kobayashi, Eduardo Iglesius, Naoya Chiba, Shizuka Shirai, Yuki Uranishi

机构 * Graduate School of Information Science and Technology, The University of Osaka(信息科学与技术研究生院,大阪大学) D3 Center, The University of Osaka(大阪大学D3中心) Graduate School of Maritime Sciences, Kobe University(海洋科学研究生院, Kobe大学)

专题命中 VLA模型 :action model(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14038 2025-07-21 cs.LG 57%

DONUT: Physics-aware Machine Learning for Real-time X-ray Nanodiffraction Analysis

Aileen Luo, Tao Zhou, Ming Du, Martin V. Holt, Andrej Singer, Mathew J. Cherukara

专题命中 VLA模型 :action model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13827 2025-07-21 cs.CL cs.IR cs.LG 57%

Question-Answer Extraction from Scientific Articles Using Knowledge Graphs and Large Language Models

Hosein Azarbonyad, Zi Long Zhu, Georgios Cheirmpos, Zubair Afzal, Vikrant Yadav, Georgios Tsatsaronis

机构 * Elsevier(埃塞伊弗)

专题命中 VLA模型 :action model(abstract);分类 cs.LG

Comments SIGIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13768 2025-07-21 cs.AI 57%

From Extraction to Synthesis: Entangled Heuristics for Agent-Augmented Strategic Reasoning

Renato Ghisellini, Remo Pareschi, Marco Pedroni, Giovanni Battista Raggi

机构 * Institute for Generative Strategy(生成战略研究所) Stake Lab, University of Molise(莫利塞大学Stake实验室)

专题命中 VLA模型 :action model(abstract);分类 cs.AI

Comments Peer-reviewed full paper accepted through a double-blind review process at the HAR 2025 conference (https://har-conf.eu/). The official version will appear in a volume of the Lecture Notes in Computer Science (LNCS) series

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14995 2025-07-17 cs.AI 57%

Learning Lifted STRIPS Models from Action Traces Alone: A Simple, General, and Scalable Solution

Jonas Gösgens, Niklas Jansen, Hector Geffner

专题命中 VLA模型 :action model(abstract);分类 cs.AI

Comments accepted at ICAPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09446 2025-07-15 cs.CV 57%

Efficient Multi-Person Motion Prediction by Lightweight Spatial and Temporal Interactions

Yuanhong Zheng, Ruixuan Yu, Jian Sun

机构 * Shandong University(山东大学) Xi’an Jiaotong University(西安交通大学) Peking University(北京大学) Pazhou Laboratory (Huangpu)(黄埔实验室)

专题命中 VLA模型 :action model(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03895 2025-07-08 cs.IR cs.AI 57%

TayFCS: Towards Light Feature Combination Selection for Deep Recommender Systems

Xianquan Wang, Zhaocheng Du, Jieming Zhu, Chuhan Wu, Qinglin Jia, Zhenhua Dong

机构 * University of Science and Technology of China(中国科学技术大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 VLA模型 :action model(abstract);分类 cs.AI

Journal ref KDD'2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23739 2025-07-01 cs.RO cs.CE cs.HC 57%

Validation of AI-Based 3D Human Pose Estimation in a Cyber-Physical Environment

Lisa Marie Otto, Michael Kaiser, Daniel Seebacher, Steffen Müller

专题命中 VLA模型 :action model(abstract);分类 cs.RO

Comments 6 pages, 5 figures, Preprint for 2025 IEEE IAVVC (International Automated Vehicle Validation Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23196 2025-07-01 cs.CV 57%

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding

Mona Ahmadian, Amir Shirian, Frank Guerin, Andrew Gilbert

机构 * University of Surrey(塞拉利昂大学)

专题命中 VLA模型 :action model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11069 2025-06-24 cs.AI cs.HC 57%

API Agents vs. GUI Agents: Divergence and Convergence

Chaoyun Zhang, Shilin He, Liqun Li, Si Qin, Yu Kang, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang

机构 * Microsoft(微软)

专题命中 VLA模型 :action model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11591 2025-06-24 math.NA cs.LG cs.NA math.DS 57%

A physics-informed neural network method for the approximation of slow invariant manifolds for the general class of stiff systems of ODEs

Dimitrios G. Patsatzis, Lucia Russo, Constantinos Siettos

机构 * Modelling Engineering Risk and Complexity, Scuola Superiore Meridionale(工程风险与复杂性建模,南方高级学校) Institute of Science and Technology for Energy and Sustainable Mobility, Consiglio Nazionale delle Ricerche(能源与可持续移动科技研究所,国家研究理事会)

专题命中 VLA模型 :action model(abstract);分类 cs.LG

Journal ref SIAM Journal on Applied Dynamical Systems 23, no. 4 (2024): 3077-3122

详情

展开后加载摘要…

URL PDF HTML 收藏