arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2773 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2773 篇

2306.11335 2024-03-21 cs.RO cs.AI cs.CV cs.LG 62%

Surfer: Progressive Reasoning with World Models for Robotic Manipulation

Pengzhen Ren, Kaidong Zhang, Hetao Zheng, Zixuan Li, Yuhang Wen, Fengda Zhu, Mas Ma, Xiaodan Liang

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.17785 2024-03-08 cs.SD cs.AI eess.AS 62%

ByteComposer: a Human-like Melody Composition Method based on Language Model Agent

Xia Liang, Xingjian Du, Jiaju Lin, Pei Zou, Yuan Wan, Bilei Zhu

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.17930 2024-02-29 cs.AI cs.CL cs.LG 62%

Pragmatic Instruction Following and Goal Assistance via Cooperative Language-Guided Inverse Planning

Tan Zhi-Xuan, Lance Ying, Vikash Mansinghka, Joshua B. Tenenbaum

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

Comments Accepted to AAMAS 2024. 8 pages (excl. references), 5 figures/tables. (Appendix: 8 pages, 8 figures/tables). Code available at: https://github.com/probcomp/CLIPS.jl

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.03047 2024-01-23 cs.CV cs.CL cs.RO 62%

ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments

Dong An, Hanqing Wang, Wenguan Wang, Zun Wang, Yan Huang, Keji He, Liang Wang

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV、cs.CL

Comments Project page: https://github.com/MarSaKi/ETPNav

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07899 2024-01-17 q-bio.QM cs.AI cs.CV cs.LG 62%

Morphological Profiling for Drug Discovery in the Era of Deep Learning

Qiaosi Tang, Ranjala Ratnayake, Gustavo Seabra, Zhe Jiang, Ruogu Fang, Lina Cui, Yousong Ding, Tamer Kahveci, Jiang Bian, Chenglong Li, Hendrik Luesch, Yanjun Li

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

Comments 44 pages, 5 figure, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.12344 2023-10-20 cs.CL cs.CV 62%

LACMA: Language-Aligning Contrastive Learning with Meta-Actions for Embodied Instruction Following

Cheng-Fu Yang, Yen-Chun Chen, Jianwei Yang, Xiyang Dai, Lu Yuan, Yu-Chiang Frank Wang, Kai-Wei Chang

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.CL

Comments EMNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.16534 2023-09-29 cs.CV cs.AI cs.LG cs.RO 62%

MotionLM: Multi-Agent Motion Forecasting as Language Modeling

Ari Seff, Brian Cera, Dian Chen, Mason Ng, Aurick Zhou, Nigamaa Nayakanti, Khaled S. Refaat, Rami Al-Rfou, Benjamin Sapp

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

Comments To appear at the International Conference on Computer Vision (ICCV) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11307 2023-09-21 cs.CL cs.AI 62%

Rating Prediction in Conversational Task Assistants with Behavioral and Conversational-Flow Features

Rafael Ferreira, David Semedo, João Magalhães

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.15021 2023-09-15 cs.RO cs.AI cs.CV cs.LG 62%

EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Yao Mu, Qinglong Zhang, Mengkang Hu, Wenhai Wang, Mingyu Ding, Jun Jin, Bin Wang, Jifeng Dai, Yu Qiao, Ping Luo

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.15097 2023-08-30 cs.AI cs.CL 62%

Sequential annotations for naturally-occurring HRI: first insights

Lucien Tisserand, Frédéric Armetta, Heike Baldauf-Quilliatre, Antoine Bouquin, Salima Hassas, Mathieu Lefort

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

Comments Peer-reviewed workshop paper accepted for the ''Human-Robot Conversational Interaction'' workshop that took place at the ''ACM/IEEE International Conference on Human-Robot Interaction'' 2023 Conference in Stockholm, Sweden

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11877 2023-08-25 cs.CV cs.AI 62%

Integrated Image and Location Analysis for Wound Classification: A Deep Learning Approach

Yash Patel, Tirth Shah, Mrinal Kanti Dhar, Taiyu Zhang, Jeffrey Niezgoda, Sandeep Gopalakrishnan, Zeyun Yu

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.16207 2023-06-29 cs.AI cs.CL cs.RO 62%

Inferring the Goals of Communicating Agents from Actions and Instructions

Lance Ying, Tan Zhi-Xuan, Vikash Mansinghka, Joshua B. Tenenbaum

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CL、cs.AI

Comments 8 pages, 5 figures. Accepted to the ICML 2023 Workshop on Theory of Mind in Communicating Agents. Supplementary Information: https://osf.io/gh758/

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.14911 2023-06-28 cs.CL cs.AI 62%

"You might think about slightly revising the title": identifying hedges in peer-tutoring interactions

Yann Raphalen, Chloé Clavel, Justine Cassell

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

Comments Published in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), 2022

Journal ref Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), Volume 1: long papers (2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.09349 2023-06-13 cs.AI cs.CL cs.RO 62%

LLM as A Robotic Brain: Unifying Egocentric Memory and Control

Jinjie Mai, Jun Chen, Bing Li, Guocheng Qian, Mohamed Elhoseiny, Bernard Ghanem

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

Comments This early project is now integrated to: Mindstorms in Natural Language-Based Societies of Mind, arXiv:2305.17066

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.06358 2023-05-12 cs.AI cs.CL 62%

Accessible Instruction-Following Agent

Kairui Zhou

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.09448 2023-04-20 cs.LG cs.CL cs.CV 62%

EC^2: Emergent Communication for Embodied Control

Yao Mu, Shunyu Yao, Mingyu Ding, Ping Luo, Chuang Gan

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.CL

Comments Published in CVPR2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.09474 2023-04-05 cs.CV cs.AI cs.RO 62%

3D Object Detection for Autonomous Driving: A Comprehensive Survey

Jiageng Mao, Shaoshuai Shi, Xiaogang Wang, Hongsheng Li

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted to International Journal of Computer Vision (IJCV). Project page is at https://github.com/PointsCoder/Awesome-3D-Object-Detection-for-Autonomous-Driving

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.08620 2023-03-27 cs.CL cs.AI cs.CY cs.HC cs.LG 62%

POTATO: The Portable Text Annotation Tool

Jiaxin Pei, Aparna Ananthasubramaniam, Xingyao Wang, Naitian Zhou, Jackson Sargent, Apostolos Dedeloudis, David Jurgens

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2022 DEMO

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.15027 2023-01-18 cs.LG cs.AI cs.CL cs.NE 62%

Symbol Emergence as Inter-personal Categorization with Head-to-head Latent Word

Kazuma Furukawa, Akira Taniguchi, Yoshinobu Hagiwara, Tadahiro Taniguchi

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

Comments 7 pages, 4 figures, 5 tables

Journal ref IEEE International Conference on Development and Learning (ICDL 2022), 2022, 60-67

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.00433 2023-01-03 cs.AI cs.CV cs.IT math.IT 62%

Optimization of Image Transmission in a Cooperative Semantic Communication Networks

Wenjing Zhang, Yining Wang, Mingzhe Chen, Tao Luo, Dusit Niyato

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

Comments 29 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.08729 2022-12-20 cs.RO cs.AI cs.CV cs.LG cs.SY eess.SY 62%

Distribution-aware Goal Prediction and Conformant Model-based Planning for Safe Autonomous Driving

Jonathan Francis, Bingqing Chen, Weiran Yao, Eric Nyberg, Jean Oh

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted: 1st Workshop on Safe Learning for Autonomous Driving, at the International Conference on Machine Learning (ICML 2022); Best Paper Award

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.06175 2022-11-14 cs.AI cs.CL cs.LG cs.RO 62%

A Generalist Agent

Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, Tom Eccles, Jake Bruce, Ali Razavi, Ashley Edwards, Nicolas Heess, Yutian Chen, Raia Hadsell, Oriol Vinyals, Mahyar Bordbar, Nando de Freitas

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Published at TMLR, 42 pages

Journal ref Transactions on Machine Learning Research, 11/2022, https://openreview.net/forum?id=1ikK0kHjvj

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06155 2022-10-17 cs.CL cs.AI 62%

ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding

Qiming Peng, Yinxu Pan, Wenjin Wang, Bin Luo, Zhenyu Zhang, Zhengjie Huang, Teng Hu, Weichong Yin, Yongfeng Chen, Yin Zhang, Shikun Feng, Yu Sun, Hao Tian, Hua Wu, Haifeng Wang

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2022 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.02764 2022-03-08 cs.CV cs.CL cs.RO 62%

Bridging the Gap Between Learning in Discrete and Continuous Environments for Vision-and-Language Navigation

Yicong Hong, Zun Wang, Qi Wu, Stephen Gould

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.02173 2022-03-04 cs.RO cs.AI cs.CV cs.LG cs.MA 62%

Multi-Agent Variational Occlusion Inference Using People as Sensors

Masha Itkina, Ye-Ji Mun, Katherine Driggs-Campbell, Mykel J. Kochenderfer

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

Comments 12 pages, 9 figures, International Conference on Robotics and Automation (ICRA) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.11576 2021-11-26 cs.LG cs.CL cs.CV 62%

Building Goal-Oriented Dialogue Systems with Situated Visual Context

Sanchit Agarwal, Jan Jezabek, Arijit Biswas, Emre Barut, Shuyang Gao, Tagyoung Chung

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.00956 2021-09-02 cs.LG cs.AI cs.CL 62%

SocialAI: Benchmarking Socio-Cognitive Abilities in Deep Reinforcement Learning Agents

Grgur Kovač, Rémy Portelas, Katja Hofmann, Pierre-Yves Oudeyer

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

Comments under review. This paper extends and generalizes work in arXiv:2104.13207

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.02846 2021-08-09 cs.AI cs.CV cs.HC cs.LG cs.RO 62%

Communicative Learning with Natural Gestures for Embodied Navigation Agents with Human-in-the-Scene

Qi Wu, Cheng-Ju Wu, Yixin Zhu, Jungseock Joo

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

Comments To appear in IROS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.13073 2021-06-24 cs.CL cs.AI 62%

Maria: A Visual Experience Powered Conversational Agent

Zujie Liang, Huang Hu, Can Xu, Chongyang Tao, Xiubo Geng, Yining Chen, Fan Liang, Daxin Jiang

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

Comments Accepted by ACL 2021 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.10110 2021-06-21 cs.CV cs.AI cs.MA cs.RO 62%

Towards Distraction-Robust Active Visual Tracking

Fangwei Zhong, Peng Sun, Wenhan Luo, Tingyun Yan, Yizhou Wang

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV、cs.AI

Comments To appear in ICML2021

详情

展开后加载摘要…

URL PDF HTML 收藏