arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

共收录 9819 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 9146 篇

2504.14618 2025-04-22 cs.CV cs.AI 62%

VM-BHINet:Vision Mamba Bimanual Hand Interaction Network for 3D Interacting Hand Mesh Recovery From a Single RGB Image

Han Bi, Ge Yu, Yu He, Wenzhuo Liu, Zijie Zheng

专题命中 VLA模型 :action model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.16882 2025-02-21 physics.chem-ph cs.AI cs.LG q-bio.BM 62%

Revealing the Relationship Between Publication Bias and Chemical Reactivity with Contrastive Learning

Wenhao Gao, Priyanka Raghavan, Ron Shprints, Connor W. Coley

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09304 2025-01-17 cs.CV cs.LG 62%

Finding the Trigger: Causal Abductive Reasoning on Video Events

Thao Minh Le, Vuong Le, Kien Do, Sunil Gupta, Svetha Venkatesh, Truyen Tran

专题命中 VLA模型 :action model(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00065 2025-01-03 cs.LG cs.AI 62%

Predicting Preschoolers' Externalizing Problems with Mother-Child Interaction Dynamics and Deep Learning

Xi Chen, Yu Ji, Cong Xia, Wen Wu

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

Comments 34 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08109 2025-01-03 cs.RO cs.CV 62%

VANP: Learning Where to See for Navigation with Self-Supervised Vision-Action Pre-Training

Mohammad Nazeri, Junzhe Wang, Amirreza Payandeh, Xuesu Xiao

专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.CV

Comments Extended version of the paper accepted at IROS 2024. Code: https://github.com/mhnazeri/VANP

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08618 2024-12-12 cs.CV cs.AI 62%

Image Retrieval Methods in the Dissimilarity Space

Madhu Kiran, Kartikey Vishnu, Rafael M. O. Cruz, Eric Granger

专题命中 VLA模型 :action model(abstract);分类 cs.CV、cs.AI

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01537 2024-12-02 cs.CV cs.RO 62%

SceneMotion: From Agent-Centric Embeddings to Scene-Wide Forecasts

Royden Wagner, Ömer Sahin Tas, Marlon Steiner, Fabian Konstantinidis, Hendrik Königshof, Marvin Klemp, Carlos Fernandez, Christoph Stiller

专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.CV

Comments ITSC'24; updated table VI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.15506 2024-11-12 cs.AI cs.CL cs.LG 62%

AgentOhana: Design Unified Data and Training Pipeline for Effective Agent Learning

Jianguo Zhang, Tian Lan, Rithesh Murthy, Zhiwei Liu, Weiran Yao, Ming Zhu, Juntao Tan, Thai Hoang, Zuxin Liu, Liangwei Yang, Yihao Feng, Shirley Kokane, Tulika Awalgaonkar, Juan Carlos Niebles, Silvio Savarese, Shelby Heinecke, Huan Wang, Caiming Xiong

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

Comments Add GitHub repo link at \url{https://github.com/SalesforceAIResearch/xLAM} and HuggingFace model link at \url{https://huggingface.co/Salesforce/xLAM-v0.1-r}

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.01819 2024-10-29 cs.AI cs.HC cs.LG 62%

A Model for Intelligible Interaction Between Agents That Predict and Explain

A. Baskar, Ashwin Srinivasan, Michael Bain, Enrico Coiera

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

Comments arXiv admin note: text overlap with arXiv:2205.08954

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.18681 2024-10-08 cs.LG cs.AI 62%

Deep Fusion: Capturing Dependencies in Contrastive Learning via Transformer Projection Heads

Huanran Li, Daniel Pimentel-Alarcón

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01571 2024-10-02 cs.CV cs.LG 62%

Counterfactual Explanations for Medical Image Classification and Regression using Diffusion Autoencoder

Matan Atad, David Schinz, Hendrik Moeller, Robert Graf, Benedikt Wiestler, Daniel Rueckert, Nassir Navab, Jan S. Kirschke, Matthias Keicher

专题命中 VLA模型 :action model(abstract);分类 cs.CV、cs.LG

Comments Accepted for publication at the Journal of Machine Learning for Biomedical Imaging (MELBA) https://melba-journal.org/2024:024. arXiv admin note: text overlap with arXiv:2303.12031

Journal ref Machine.Learning.for.Biomedical.Imaging. 2 (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13951 2024-09-24 cs.CV cs.LG eess.IV 62%

Deep learning for fast segmentation and critical dimension metrology & characterization enabling AR/VR design and fabrication

Kundan Chaudhary, Subhei Shaar, Raja Muthinti

专题命中 VLA模型 :action model(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08853 2024-09-16 cs.AI cs.RO 62%

Using The Concept Hierarchy for Household Action Recognition

Andrei Costinescu, Luis Figueredo, Darius Burschka

专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.AI

Comments 5 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.12406 2024-08-06 cs.AI cs.LG 62%

Offline Imitation of Badminton Player Behavior via Experiential Contexts and Brownian Motion

Kuang-Da Wang, Wei-Yao Wang, Ping-Chun Hsieh, Wen-Chih Peng

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

Comments Accepted by the European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19510 2024-07-30 cs.RO cs.CV 62%

EPD: Long-term Memory Extraction, Context-awared Planning and Multi-iteration Decision @ EgoPlan Challenge ICML 2024

Letian Shi, Qi Lv, Xiang Deng, Liqiang Nie

专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06549 2024-07-10 cs.IR cs.AI cs.CL cs.LG 62%

AutoTask: Task Aware Multi-Faceted Single Model for Multi-Task Ads Relevance

Shouchang Guo, Sonam Damani, Keng-hao Chang

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17968 2024-06-27 cs.IR cs.AI cs.LG stat.ML 62%

Efficient Document Ranking with Learnable Late Interactions

Ziwei Ji, Himanshu Jain, Andreas Veit, Sashank J. Reddi, Sadeep Jayasumana, Ankit Singh Rawat, Aditya Krishna Menon, Felix Yu, Sanjiv Kumar

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04043 2024-06-10 cs.CV cs.AI 62%

Doodle Your 3D: From Abstract Freehand Sketches to Precise 3D Shapes

Hmrishav Bandyopadhyay, Subhadeep Koley, Ayan Das, Ayan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury, Tao Xiang, Yi-Zhe Song

专题命中 VLA模型 :action model(abstract);分类 cs.CV、cs.AI

Comments CVPR 2024, Project Page: https://hmrishavbandy.github.io/doodle23d/

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05452 2024-05-02 astro-ph.IM cs.CV cs.LG 62%

The R2D2 deep neural network series paradigm for fast precision imaging in radio astronomy

Amir Aghabiglou, Chung San Chu, Arwa Dabbech, Yves Wiaux

专题命中 VLA模型 :VLA(abstract);分类 cs.CV、cs.LG

Comments Accepted for publication in ApJS

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.06234 2024-04-02 cs.RO cs.LG cs.SY eess.SY 62%

EVORA: Deep Evidential Traversability Learning for Risk-Aware Off-Road Autonomy

Xiaoyi Cai, Siddharth Ancha, Lakshay Sharma, Philip R. Osteen, Bernadette Bucher, Stephen Phillips, Jiuguang Wang, Michael Everett, Nicholas Roy, Jonathan P. How

专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.LG

Comments Under review. Journal extension for arXiv:2210.00153. Project website: https://xiaoyi-cai.github.io/evora/

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.08978 2024-03-19 cs.RO cs.LG q-bio.QM 62%

Quantifying the biomimicry gap in biohybrid robot-fish pairs

Vaios Papaspyros, Guy Theraulaz, Clément Sire, Francesco Mondada

专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14486 2024-02-23 cs.GT cs.AI cs.LG econ.TH 62%

Are Bounded Contracts Learnable and Approximately Optimal?

Yurong Chen, Zhaohua Chen, Xiaotie Deng, Zhiyi Huang

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.03755 2023-12-08 cs.CL cs.AI cs.CY cs.LG 62%

Near-real-time Earthquake-induced Fatality Estimation using Crowdsourced Data and Large-Language Models

Chenguang Wang, Davis Engler, Xuechun Li, James Hou, David J. Wald, Kishor Jaiswal, Susu Xu

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.04374 2023-10-24 eess.SY cs.LG cs.RO cs.SY 62%

Learning to Identify Graphs from Node Trajectories in Multi-Robot Networks

Eduardo Sebastian, Thai Duong, Nikolay Atanasov, Eduardo Montijano, Carlos Sagues

专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.LG

Comments Accepted at IEEE MRS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04424 2023-10-10 cs.NE cs.AI cs.LG q-bio.MN 62%

Stability Analysis of Non-Linear Classifiers using Gene Regulatory Neural Network for Biological AI

Adrian Ratwatte, Samitha Somathilaka, Sasitharan Balasubramaniam, Assaf A. Gilad

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.16898 2023-10-02 cs.RO cs.CL cs.CV cs.HC 62%

A Sign Language Recognition System with Pepper, Lightweight-Transformer, and LLM

JongYoon Lim, Inkyu Sa, Bruce MacDonald, Ho Seok Ahn

专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05787 2023-09-13 cs.AI cs.HC cs.LG 62%

Adaptive User-centered Neuro-symbolic Learning for Multimodal Interaction with Autonomous Systems

Amr Gomaa, Michael Feld

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

Comments AI&HCI Workshop accepted paper at ICML2023 and accepted at ICMI2023 Blue Sky Papers. arXiv admin note: text overlap with arXiv:2211.03539

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12537 2023-08-25 cs.RO cs.CV 62%

HuBo-VLM: Unified Vision-Language Model designed for HUman roBOt interaction tasks

Zichao Dong, Weikun Zhang, Xufeng Huang, Hang Ji, Xin Zhan, Junbo Chen

专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.03629 2023-08-09 cs.CL cs.AI cs.LG 62%

MedMine: Examining Pre-trained Language Models on Medication Mining

Haifa Alrdahi, Lifeng Han, Hendrik Šuvalov, Goran Nenadic

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

Comments Open Research Project. 7 pages, 1 figure, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.02723 2023-08-08 cs.SD cs.AI cs.LG cs.MM eess.AS 62%

Towards Improving Harmonic Sensitivity and Prediction Stability for Singing Melody Extraction

Keren Shao, Ke Chen, Taylor Berg-Kirkpatrick, Shlomo Dubnov

专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG

Comments 7 pages, 4 figures, 2 tables, Proceedings of the 24th International Society for Music Information Retrieval Conference, ISMIR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏