arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2777 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2777 篇

2401.12624 2024-03-05 cs.AI cs.IT cs.LG cs.NI math.IT 57%

Knowledge Distillation from Language-Oriented to Emergent Communication for Multi-Agent Remote Control

Yongjun Kim, Sejin Seo, Jihong Park, Mehdi Bennis, Seong-Lyun Kim, Junil Choi

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.07998 2024-02-13 cs.CL 57%

What Artificial Neural Networks Can Tell Us About Human Language Acquisition

Alex Warstadt, Samuel R. Bowman

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

Comments Please cite the published version with the following information: @incollection{warstadt2022artificial, title={What artificial neural networks can tell us about human language acquisition}, author={Warstadt, Alex and Bowman, Samuel R.}, booktitle={Algebraic Structures in Natural Language}, pages={17--60}, year={2022}, publisher={CRC Press} }

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.05440 2024-02-09 cs.CL 57%

Improving Agent Interactions in Virtual Environments with Language Models

Jack Zhang

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.11459 2024-01-23 cs.AR cs.AI cs.LG 57%

AttentionLego: An Open-Source Building Block For Spatially-Scalable Large Language Model Accelerator With Processing-In-Memory Technology

Rongqing Cong, Wenyang He, Mingxuan Li, Bangning Luo, Zebin Yang, Yuchao Yang, Ru Huang, Bonan Yan

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments for associated source codes, see https://bonany.cc/attentionleg

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.04334 2024-01-10 cs.RO cs.AI 57%

Large Language Models for Robotics: Opportunities, Challenges, and Perspectives

Jiaqi Wang, Zihao Wu, Yiwei Li, Hanqi Jiang, Peng Shu, Enze Shi, Huawen Hu, Chong Ma, Yiheng Liu, Xuhui Wang, Yincheng Yao, Xuan Liu, Huaqin Zhao, Zhengliang Liu, Haixing Dai, Lin Zhao, Bao Ge, Xiang Li, Tianming Liu, Shu Zhang

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11561 2023-12-15 cs.CV 57%

Target-Grounded Graph-Aware Transformer for Aerial Vision-and-Dialog Navigation

Yifei Su, Dong An, Yuan Xu, Kehan Chen, Yan Huang

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV

Comments 1st Place Solution for the AVDN Challenge in ICCV CLVL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06037 2023-12-13 cs.AI 57%

Multimodality of AI for Education: Towards Artificial General Intelligence

Gyeong-Geon Lee, Lehong Shi, Ehsan Latif, Yizhu Gao, Arne Bewersdorff, Matthew Nyaaba, Shuchen Guo, Zihao Wu, Zhengliang Liu, Hui Wang, Gengchen Mai, Tiaming Liu, Xiaoming Zhai

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.00735 2023-12-12 cs.RO cs.AI cs.LG 57%

MTP-GO: Graph-Based Probabilistic Multi-Agent Trajectory Prediction with Neural ODEs

Theodor Westny, Joel Oskarsson, Björn Olofsson, Erik Frisk

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments Code: https://github.com/westny/mtp-go

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13255 2023-12-08 cs.CV 57%

Steve-Eye: Equipping LLM-based Embodied Agents with Visual Perception in Open Worlds

Sipeng Zheng, Jiazheng Liu, Yicheng Feng, Zongqing Lu

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments 19 pages, 19 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.18553 2023-12-01 cs.LG cs.CV cs.RO 57%

Heterogeneous Graph-based Trajectory Prediction using Local Map Context and Social Interactions

Daniel Grimm, Maximilian Zipfl, Felix Hertlein, Alexander Naumann, Jürgen Lüttin, Steffen Thoma, Stefan Schmid, Lavdim Halilaj, Achim Rettinger, J. Marius Zöllner

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

Comments Accepted on IEEE ITSC 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.18307 2023-12-01 cs.LG cs.CV cs.RO 57%

Categorical Traffic Transformer: Interpretable and Diverse Behavior Prediction with Tokenized Latent

Yuxiao Chen, Sander Tonkens, Marco Pavone

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.15033 2023-11-28 cs.RO cs.AI 57%

Agent as Cerebrum, Controller as Cerebellum: Implementing an Embodied LMM-based Agent on Drones

Haoran Zhao, Fengxing Pan, Huqiuyue Ping, Yaoming Zhou

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments 17 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.13169 2023-11-23 cs.LG cs.AI 57%

SiGeo: Sub-One-Shot NAS via Information Theory and Geometry of Loss Landscape

Hua Zheng, Kuang-Hung Liu, Igor Fedorov, Xin Zhang, Wen-Yen Chen, Wei Wen

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments 24 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.11184 2023-11-21 cs.CV 57%

Diverse Shape Completion via Style Modulated Generative Adversarial Networks

Wesley Khademi, Li Fuxin

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02933 2023-11-15 cs.HC cs.AI cs.RO 57%

In Time and Space: Towards Usable Adaptive Control for Assistive Robotic Arms

Max Pascher, Kirill Kronhardt, Felix Ferdinand Goldau, Udo Frese, Jens Gerken

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments RO-MAN'23: 32nd IEEE International Conference on Robot and Human Interactive Communication, Busan, Korea

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12736 2023-11-08 cs.CV 57%

FastSurfer-HypVINN: Automated sub-segmentation of the hypothalamus and adjacent structures on high-resolutional brain MRI

Santiago Estrada, David Kügler, Emad Bahrami, Peng Xu, Dilshad Mousa, Monique M. B. Breteler, N. Ahmad Aziz, Martin Reuter

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

Comments Accepted in Imaging Neuroscience

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.06775 2023-11-02 cs.HC cs.AI 57%

Conceptual Framework for Autonomous Cognitive Entities

David Shapiro, Wangfan Li, Manuel Delaflor, Carlos Toxtli

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments 34 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.19240 2023-10-31 cs.LG cs.AI 57%

NetHack is Hard to Hack

Ulyana Piterbarg, Lerrel Pinto, Rob Fergus

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10095 2023-10-17 eess.IV cs.CV cs.LG 57%

A Multi-Scale Spatial Transformer U-Net for Simultaneously Automatic Reorientation and Segmentation of 3D Nuclear Cardiac Images

Yangfan Ni, Duo Zhang, Gege Ma, Lijun Lu, Zhongke Huang, Wentao Zhu

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments 17 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.10842 2023-09-28 cs.AI cs.LG 57%

Enhancing Agent Communication and Learning through Action and Language

Hugo Caselles-Dupré, Olivier Sigaud, Mohamed Chetouani

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments IMOL workshop, Paris 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11283 2023-09-21 cs.CL 57%

The Wizard of Curiosities: Enriching Dialogues with Fun Facts

Frederico Vicente, Rafael Ferreira, David Semedo, João Magalhães

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10375 2023-09-20 cs.CV 57%

Pointing out Human Answer Mistakes in a Goal-Oriented Visual Dialogue

Ryosuke Oshima, Seitaro Shinagawa, Hideki Tsunashima, Qi Feng, Shigeo Morishima

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments Accepted at ICCVW 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.00923 2023-09-19 cs.RO cs.CV cs.HC 57%

Sonicverse: A Multisensory Simulation Platform for Embodied Household Agents that See and Hear

Ruohan Gao, Hao Li, Gokul Dharan, Zhuzhu Wang, Chengshu Li, Fei Xia, Silvio Savarese, Li Fei-Fei, Jiajun Wu

专题命中 多模态Agent :audio-visual(abstract);分类 cs.CV

Comments In ICRA 2023. Project page: https://ai.stanford.edu/~rhgao/sonicverse/. Code: https://github.com/StanfordVL/sonicverse. Gao and Li contributed equally to this work and are in alphabetical order

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05036 2023-09-12 cs.RO cs.CV 57%

What Is Near?: Room Locality Learning for Enhanced Robot Vision-Language-Navigation in Indoor Living Environments

Muraleekrishna Gopinathan, Jumana Abu-Khalaf, David Suter, Sidike Paheding, Nathir A. Rawashdeh

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.01073 2023-09-06 cs.CV 57%

Spatial and Visual Perspective-Taking via View Rotation and Relation Reasoning for Embodied Reference Understanding

Cheng Shi, Sibei Yang

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

Comments ECCV 2022. Code: http://github.com/ChengShiest/REP-ERU

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.10324 2023-09-01 cs.RO cs.AI cs.HC 57%

HARPS: An Online POMDP Framework for Human-Assisted Robotic Planning and Sensing

Luke Burks, Hunter M. Ray, Jamison McGinley, Sousheel Vunnam, Nisar Ahmed

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments Accepted to IEEE Transactions on Robotics. 20 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09066 2023-08-21 cs.CV 57%

PatchCT: Aligning Patch Set and Label Set with Conditional Transport for Multi-Label Image Classification

Miaoge Li, Dongsheng Wang, Xinyang Liu, Zequn Zeng, Ruiying Lu, Bo Chen, Mingyuan Zhou

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV

Comments accepted by ICCV23

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.07751 2023-08-16 cs.CV 57%

CASPNet++: Joint Multi-Agent Motion Prediction

Maximilian Schäfer, Kun Zhao, Anton Kummert

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06498 2023-08-15 cs.AI cs.HC cs.RO 57%

Latent Emission-Augmented Perspective-Taking (LEAPT) for Human-Robot Interaction

Kaiqi Chen, Jing Yu Lim, Kingsley Kuan, Harold Soh

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.05221 2023-08-11 cs.HC cs.AI cs.RO 57%

Alexa, play with robot: Introducing the First Alexa Prize SimBot Challenge on Embodied AI

Hangjie Shi, Leslie Ball, Govind Thattai, Desheng Zhang, Lucy Hu, Qiaozi Gao, Suhaila Shakiah, Xiaofeng Gao, Aishwarya Padmakumar, Bofei Yang, Cadence Chung, Dinakar Guthy, Gaurav Sukhatme, Karthika Arumugam, Matthew Wen, Osman Ipek, Patrick Lange, Rohan Khanna, Shreyas Pansare, Vasu Sharma, Chao Zhang, Cris Flagg, Daniel Pressel, Lavina Vaz, Luke Dai, Prasoon Goyal, Sattvik Sahai, Shaohua Liu, Yao Lu, Anna Gottardi, Shui Hu, Yang Liu, Dilek Hakkani-Tur, Kate Bland, Heather Rocker, James Jeun, Yadunandana Rao, Michael Johnston, Akshaya Iyengar, Arindam Mandal, Prem Natarajan, Reza Ghanadan

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏