arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2766 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2766 篇

2410.05839 2024-10-10 cs.AI cs.DB 79%

Bottom-up Anytime Discovery of Generalised Multimodal Graph Patterns for Knowledge Graphs

Xander Wilcke, Rick Mourits, Auke Rijpma, Richard Zijdeman

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15243 2024-09-24 cs.AI cs.ET cs.HC 79%

MACeIP: A Multimodal Ambient Context-enriched Intelligence Platform in Smart Cities

Truong Thanh Hung Nguyen, Phuc Truong Loc Nguyen, Monica Wachowicz, Hung Cao

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments 4 pages, 6 figures, IEEE/IEIE ICCE-Asia 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08264 2024-09-17 cs.AI 79%

Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale

Rogerio Bonatti, Dan Zhao, Francesco Bonacci, Dillon Dupont, Sara Abdali, Yinheng Li, Yadong Lu, Justin Wagle, Kazuhito Koishida, Arthur Bucker, Lawrence Jang, Zack Hui

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15299 2024-08-29 q-bio.BM cs.AI cs.LG 79%

TourSynbio: A Multi-Modal Large Model and Agent Framework to Bridge Text and Protein Sequences for Protein Engineering

Yiqing Shen, Zan Chen, Michail Mamalakis, Yungeng Liu, Tianbin Li, Yanzhou Su, Junjun He, Pietro Liò, Yu Guang Wang

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11138 2024-08-22 cs.RO cs.CV 79%

Target-Oriented Object Grasping via Multimodal Human Guidance

Pengwei Xie, Siang Chen, Dingchang Hu, Yixiang Dai, Kaiqin Yang, Guijin Wang

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

Comments Accepted by ECCV 2024 Workshop on Assistive Computer Vision and Robotics (ACVR 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02394 2024-08-06 cs.CV cs.RO 79%

CMR-Agent: Learning a Cross-Modal Agent for Iterative Image-to-Point Cloud Registration

Gongxin Yao, Yixin Xuan, Xinyang Li, Yu Pan

专题命中 多模态Agent :cross-modal(title,abstract);分类 cs.CV

Comments Accepted to IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19438 2024-07-30 cs.AI cs.HC 79%

Conversational AI Multi-Agent Interoperability, Universal Open APIs for Agentic Natural Language Multimodal Communications

Diego Gosmar, Deborah A. Dahl, Emmett Coin

专题命中 多模态Agent :multimodal(title);multi-modal(abstract);分类 cs.AI

Comments 22 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.00290 2024-07-30 cs.CV 79%

MEIA: Multimodal Embodied Perception and Interaction in Unknown Environments

Yang Liu, Xinshuai Song, Kaixuan Jiang, Weixing Chen, Jingzhou Luo, Guanbin Li, Liang Lin

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

Comments Codes will be available at https://github.com/HCPLab-SYSU/Embodied_AI_Paper_List

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11333 2024-07-17 cs.RO cs.SD eess.AS 79%

Disentangled Acoustic Fields For Multimodal Physical Scene Understanding

Jie Yin, Andrew Luo, Yilun Du, Anoop Cherian, Tim K. Marks, Jonathan Le Roux, Chuang Gan

专题命中 多模态Agent :multimodal(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.10022 2024-07-16 cs.AI cond-mat.mes-hall cond-mat.mtrl-sci cond-mat.stat-mech cs.MA 79%

AtomAgents: Alloy design and discovery through physics-aware multi-modal multi-agent artificial intelligence

Alireza Ghafarollahi, Markus J. Buehler

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.12071 2024-07-15 physics.soc-ph cs.AI 79%

Activity-based and agent-based Transport model of Melbourne (AToM): an open multi-modal transport simulation model for Greater Melbourne

Afshin Jafari, Dhirendra Singh, Alan Both, Mahsa Abdollahyar, Lucy Gunn, Steve Pemberton, Billie Giles-Corti

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01824 2024-07-03 cs.HC cs.CL cs.RO 79%

Empathic Grounding: Explorations using Multimodal Interaction and Large Language Models with Conversational Agents

Mehdi Arjmand, Farnaz Nouraei, Ian Steenstra, Timothy Bickmore

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.18485 2024-07-01 q-fin.TR cs.AI 79%

A Multimodal Foundation Agent for Financial Trading: Tool-Augmented, Diversified, and Generalist

Wentao Zhang, Lingxuan Zhao, Haochong Xia, Shuo Sun, Jiaze Sun, Molei Qin, Xinyi Li, Yuqing Zhao, Yilei Zhao, Xinyu Cai, Longtao Zheng, Xinrun Wang, Bo An

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.16850 2024-06-25 cs.CV cs.RO 79%

From Perfect to Noisy World Simulation: Customizable Embodied Multi-modal Perturbations for SLAM Robustness Benchmarking

Xiaohao Xu, Tianyi Zhang, Sibo Wang, Xiang Li, Yongqi Chen, Ye Li, Bhiksha Raj, Matthew Johnson-Roberson, Xiaonan Huang

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CV

Comments 50 pages. arXiv admin note: substantial text overlap with arXiv:2402.08125

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11941 2024-06-04 cs.CL 79%

CoCo-Agent: A Comprehensive Cognitive MLLM Agent for Smartphone GUI Automation

Xinbei Ma, Zhuosheng Zhang, Hai Zhao

专题命中 多模态Agent :MLLM(title);multimodal(abstract);分类 cs.CL

Comments ACL'2024 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11308 2024-06-04 cs.AI stat.ML 79%

MCD: A Model-Agnostic Counterfactual Search Method For Multi-modal Design Modifications

Lyle Regenwetter, Yazan Abu Obaideh, Faez Ahmed

专题命中 多模态Agent :multi-modal(title);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02968 2024-05-28 cs.CV cs.LG 79%

Delving into Multi-modal Multi-task Foundation Models for Road Scene Understanding: From Learning Paradigm Perspectives

Sheng Luo, Wei Chen, Wanxin Tian, Rui Liu, Luanxuan Hou, Xiubao Zhang, Haifeng Shen, Ruiqi Wu, Shuyi Geng, Yi Zhou, Ling Shao, Yi Yang, Bojun Gao, Qun Li, Guobin Wu

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to IEEE Transactions on Intelligent Vehicles(T-IV). 24 pages, 9 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.03783 2024-05-14 cs.AI cs.RO cs.SC 79%

Scene-Driven Multimodal Knowledge Graph Construction for Embodied AI

Song Yaoxian, Sun Penglei, Liu Haoyu, Li Zhixu, Song Wei, Xiao Yanghua, Zhou Xiaofang

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16714 2024-04-01 cs.CV 79%

Embodied Multi-Modal Agent trained by an LLM from a Parallel TextWorld

Yijun Yang, Tianyi Zhou, Kanxue Li, Dapeng Tao, Lusong Li, Li Shen, Xiaodong He, Jing Jiang, Yuhui Shi

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07472 2024-03-28 cs.CV 79%

MP5: A Multi-modal Open-ended Embodied System in Minecraft via Active Perception

Yiran Qin, Enshen Zhou, Qichang Liu, Zhenfei Yin, Lu Sheng, Ruimao Zhang, Yu Qiao, Jing Shao

专题命中 多模态Agent :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments Accepted to CVPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08446 2024-03-26 cs.LG cs.AI 79%

Towards Robust Multi-Modal Reasoning via Model Selection

Xiangyan Liu, Rongxue Li, Wei Ji, Tao Lin

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.AI

Comments Accepted by ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11854 2024-02-27 cs.LG cs.AI stat.ML 79%

Multimodal Web Navigation with Instruction-Finetuned Foundation Models

Hiroki Furuta, Kuang-Huei Lee, Ofir Nachum, Yutaka Matsuo, Aleksandra Faust, Shixiang Shane Gu, Izzeddin Gur

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments Accepted to ICLR 2024. Website: https://sites.google.com/view/mm-webnav/

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11145 2024-02-20 cs.HC cs.CV cs.LG 79%

Supporting Experts with a Multimodal Machine-Learning-Based Tool for Human Behavior Analysis of Conversational Videos

Riku Arakawa, Kiyosu Maeda, Hiromu Yakura

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.15275 2024-01-30 cs.CV 79%

Dynamic Transformer Architecture for Continual Learning of Multimodal Tasks

Yuliang Cai, Mohammad Rostami

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04318 2023-12-08 cs.AI cs.LG 79%

MIMo: A Multi-Modal Infant Model for Studying Cognitive Development

Dominik Mattern, Pierre Schumacher, Francisco M. López, Marcel C. Raabe, Markus R. Ernst, Arthur Aubret, Jochen Triesch

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.AI

Comments 11 pages, 8 figures. Submitted to IEEE Transactions on Congnitive and Developmental Systems (TCDS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.15311 2023-12-08 cs.HC cs.AI cs.GR 79%

The Importance of Multimodal Emotion Conditioning and Affect Consistency for Embodied Conversational Agents

Che-Jui Chang, Samuel S. Sohn, Sen Zhang, Rajath Jayashankar, Muhammad Usman, Mubbasir Kapadia

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.05997 2023-12-01 cs.AI 79%

JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Zihao Wang, Shaofei Cai, Anji Liu, Yonggang Jin, Jinbing Hou, Bowei Zhang, Haowei Lin, Zhaofeng He, Zilong Zheng, Yaodong Yang, Xiaojian Ma, Yitao Liang

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments update project page

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.14281 2023-11-27 cs.CV 79%

Multi-modal Instance Refinement for Cross-domain Action Recognition

Yuan Qing, Naixing Wu, Shaohua Wan, Lixin Duan

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by PRCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14260 2023-10-18 cs.CL 79%

R2H: Building Multimodal Navigation Helpers that Respond to Help Requests

Yue Fan, Jing Gu, Kaizhi Zheng, Xin Eric Wang

专题命中 多模态Agent :multimodal(title);multi-modal(abstract);分类 cs.CL

Comments EMNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.13561 2023-10-03 cs.HC cs.CV 79%

Project Aria: A New Tool for Egocentric Multi-Modal AI Research

Jakob Engel, Kiran Somasundaram, Michael Goesele, Albert Sun, Alexander Gamino, Andrew Turner, Arjang Talattof, Arnie Yuan, Bilal Souti, Brighid Meredith, Cheng Peng, Chris Sweeney, Cole Wilson, Dan Barnes, Daniel DeTone, David Caruso, Derek Valleroy, Dinesh Ginjupalli, Duncan Frost, Edward Miller, Elias Mueggler, Evgeniy Oleinik, Fan Zhang, Guruprasad Somasundaram, Gustavo Solaira, Harry Lanaras, Henry Howard-Jenkins, Huixuan Tang, Hyo Jin Kim, Jaime Rivera, Ji Luo, Jing Dong, Julian Straub, Kevin Bailey, Kevin Eckenhoff, Lingni Ma, Luis Pesqueira, Mark Schwesinger, Maurizio Monge, Nan Yang, Nick Charron, Nikhil Raina, Omkar Parkhi, Peter Borschowa, Pierre Moulon, Prince Gupta, Raul Mur-Artal, Robbie Pennington, Sachin Kulkarni, Sagar Miglani, Santosh Gondi, Saransh Solanki, Sean Diener, Shangyi Cheng, Simon Green, Steve Saarinen, Suvam Patra, Tassos Mourikis, Thomas Whelan, Tripti Singh, Vasileios Balntas, Vijay Baiyya, Wilson Dreewes, Xiaqing Pan, Yang Lou, Yipu Zhao, Yusuf Mansour, Yuyang Zou, Zhaoyang Lv, Zijian Wang, Mingfei Yan, Carl Ren, Renzo De Nardi, Richard Newcombe

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏