arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2777 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2777 篇

2412.15876 2025-09-08 cs.HC cs.AI cs.GR 57%

AI-in-the-loop: The future of biomedical visual analytics applications in the era of AI

Katja Bühler, Thomas Höllt, Thomas Schulz, Pere-Pau Vázquez

机构 * Vienna Research Center for Visual Computing(维也纳视觉计算研究中心) VRVis GmbH(VRVis公司) Delft University of Technology(代尔夫特理工大学) University of Bonn(波恩大学) Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔机器学习与人工智能研究所) Universitat Politècnica de Catalunya(巴塞罗那理工大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments Accepted for publication in IEEE Computer Graphics & Applications

Journal ref K. Bühler, T. Hollt, T. Schultz and P. Vazquez, "AI-in-The-Loop: The Future of Biomedical Visual Analytics Applications in the Era of AI" in IEEE Computer Graphics and Applications, vol. 45, no. 02, pp. 90-99, March-April 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03536 2025-09-05 cs.AI cs.HC 57%

PG-Agent: An Agent Powered by Page Graph

Weizhi Chen, Ziwei Wang, Leyang Yang, Sheng Zhou, Xiaoxuan Tang, Jiajun Bu, Yong Li, Wei Jiang

机构 * Zhejiang Key Lab of Accessible Perception \& Intelligent Systems, Zhejiang University Hangzhou China Zhejiang University

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments Paper accepted to ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00218 2025-09-04 cs.RO cs.AI 57%

Embodied AI in Social Spaces: Responsible and Adaptive Robots in Complex Setting -- UKAIRS 2025 (Copy)

Aleksandra Landowska, Aislinn D Gomez Bergin, Ayodeji O. Abioye, Jayati Deshmukh, Andriana Bouadouki, Maria Wheadon, Athina Georgara, Dominic Price, Tuyen Nguyen, Shuang Ao, Lokesh Singh, Yi Long, Raffaele Miele, Joel E. Fischer, Sarvapali D. Ramchurn

机构 * School of Computer Science, University of Nottingham(诺丁汉大学计算机科学学院) School of Computing and Communications, The Open University(开放大学计算与通讯学院) School of Electronics and Computer Science, University of Southampton(南安普顿大学电子与计算机科学学院) Responsible AI, University of Southampton(南安普顿大学负责任的人工智能) University of Liverpool(利兹大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18096 2025-09-04 cs.AI 57%

Deep Research Agents: A Systematic Examination And Roadmap

Yuxuan Huang, Yihang Chen, Haozheng Zhang, Kang Li, Huichi Zhou, Meng Fang, Linyi Yang, Xiaoguang Li, Lifeng Shang, Songcen Xu, Jianye Hao, Kun Shao, Jun Wang

机构 * Deep Research Agents(深度研究代理)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02322 2025-09-03 cs.CV 57%

OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds

Longrong Yang, Zhixiong Zeng, Yufeng Zhong, Jing Huang, Liming Zheng, Lei Chen, Haibo Qiu, Zequn Qin, Lin Ma, Xi Li

机构 * Meituan(美团) Zhejiang University(浙江大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03735 2025-09-03 cs.CV 57%

Multi-Agent System for Comprehensive Soccer Understanding

Jiayuan Rao, Zifeng Li, Haoning Wu, Ya Zhang, Yanfeng Wang, Weidi Xie

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments Accepted by ACM MM 2025; Project Page: https://jyrao.github.io/SoccerAgent/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16271 2025-09-03 cs.CV cs.LG 57%

Structuring GUI Elements through Vision Language Models: Towards Action Space Generation

Yi Xu, Yesheng Zhang, Jiajia Liu, Jingdong Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Ant Group(蚂蚁集团)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments 10pageV0

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21407 2025-09-03 cs.AI 57%

Graph-Augmented Large Language Model Agents: Current Progress and Future Prospects

Yixin Liu, Guibin Zhang, Kun Wang, Shiyuan Li, Shirui Pan

机构 * Griffith University(格里菲斯大学) National University of Singapore(国立新加坡大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments 15 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20370 2025-08-29 cs.SE cs.AI 57%

Adaptive Root Cause Localization for Microservice Systems with Multi-Agent Recursion-of-Thought

Lingzhe Zhang, Tong Jia, Kangjin Wang, Weijie Hong, Chiming Duan, Minghua He, Ying Li

机构 * Peking University(北京大学) Alibaba Group(阿里巴巴集团)

专题命中 多模态Agent :cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04447 2025-08-27 cs.CV cs.RO 57%

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

Wenyao Zhang, Hongsi Liu, Zekun Qi, Yunnan Wang, Xinqiang Yu, Jiazhao Zhang, Runpei Dong, Jiawei He, Fan Lu, He Wang, Zhizheng Zhang, Li Yi, Wenjun Zeng, Xin Jin

机构 * SJTU(上海交通大学) EIT(欧洲研究所) THU(清华大学) Galbot PKU(北京大学) UIUC(伊利诺伊大学香槟分校) USTC(中国科学技术大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17380 2025-08-26 cs.AI 57%

Mimicking the Physicist's Eye:A VLM-centric Approach for Physics Formula Discovery

Jiaqi Liu, Songning Lai, Pengze Li, Di Yu, Wenjie Zhou, Yiyang Zhou, Peng Xia, Zijun Wang, Xi Chen, Shixiang Tang, Lei Bai, Wanli Ouyang, Mingyu Ding, Huaxiu Yao, Aoran Wang

机构 * UNC–Chapel Hill(北卡罗来纳大学教堂山分校) HKUST (Guangzhou)(香港科技大学(广州)) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Fudan University(复旦大学) Tsinghua University(清华大学) Nankai University(南开大学) UC Santa Cruz(圣塔克鲁兹大学) The Chinese University of Hong Kong(香港中文大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17198 2025-08-26 cs.AI 57%

From reactive to cognitive: brain-inspired spatial intelligence for embodied agents

Shouwei Ruan, Liyuan Wang, Caixin Kang, Qihui Zhu, Songming Liu, Xingxing Wei, Hang Su

机构 * Department of Computer Science and Technology, Institute for AI, BNRist Center, Tsinghua-Bosch Joint ML Center, THBI Lab(计算机科学与技术系、人工智能研究院、BNRist中心、清华-博世联合机器学习中心、THBI实验室) Institute of Artificial Intelligence(人工智能研究院) Department of Psychological and Cognitive Sciences(心理学与认知科学系)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments 40 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15164 2025-08-22 cs.CL 57%

ContextualLVLM-Agent: A Holistic Framework for Multi-Turn Visually-Grounded Dialogue and Complex Instruction Following

Seungmin Han, Haeun Kwon, Ji-jun Park, Taeyang Yoon

机构 * Dongguk University(东国大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21969 2025-08-22 cs.RO cs.AI 57%

Embodied Long Horizon Manipulation with Closed-loop Code Generation and Incremental Few-shot Adaptation

Yuan Meng, Xiangtong Yao, Haihui Ye, Yirui Zhou, Shengqiang Zhang, Zhenguo Sun, Xukun Li, Zhenshan Bing, Alois Knoll

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments update ICRA 6 page

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10833 2025-08-18 cs.CV 57%

UI-Venus Technical Report: Building High-performance UI Agents with RFT

Zhangxuan Gu, Zhengwen Zeng, Zhenyu Xu, Xingran Zhou, Shuheng Shen, Yunfei Liu, Beitong Zhou, Changhua Meng, Tianyu Xia, Weizhi Chen, Yue Wen, Jingya Dou, Fei Tang, Jinzhen Lin, Yulin Liu, Zhenlin Guo, Yichen Gong, Heng Jia, Changlong Gao, Yuan Guo, Yong Deng, Zhenyu Guo, Liang Chen, Weiqiang Wang

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17593 2025-08-18 cs.HC cs.AI 57%

JELAI: Integrating AI and Learning Analytics in Jupyter Notebooks

Manuel Valle Torre, Thom van der Velden, Marcus Specht, Catharine Oertel

机构 * Delft University of Technology(代尔夫特理工大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments Accepted for AIED 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09858 2025-08-14 cs.CV 57%

HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics

Weiqi Li, Zehao Zhang, Liang Lin, Guangrun Wang

专题命中 多模态Agent :MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09318 2025-08-14 cs.LO cs.AI 57%

TPTP World Infrastructure for Non-classical Logics

Alexander Steen, Geoff Sutcliffe

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments 35 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07466 2025-08-12 cs.AI 57%

Grounding Natural Language for Multi-agent Decision-Making with Multi-agentic LLMs

Dom Huh, Prasant Mohapatra

机构 * UC Davis(加州大学戴维斯分校) University of South Florida(佛罗里达州立大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07010 2025-08-12 cs.MM cs.HC cs.MA 57%

Narrative Memory in Machines: Multi-Agent Arc Extraction in Serialized TV

Roberto Balestri, Guglielmo Pescatore

专题命中 多模态Agent :multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06342 2025-08-11 cs.CV cs.SI 57%

Street View Sociability: Interpretable Analysis of Urban Social Behavior Across 15 Cities

Kieran Elrod, Katherine Flanigan, Mario Bergés

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00342 2025-08-11 cs.CV 57%

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering

Zechuan Li, Hongshan Yu, Yihao Ding, Yan Li, Yong He, Naveed Akhtar

机构 * organization= College of Electrical Information Engineering,Hunan University , city= Changsha , postcode= 410082 , state= Hunan , country= China organization= School of Computing \& Information Systems ,The University of Melbourne , city= Melbourne , postcode= VIC 3053 , state= VIC , country= Australia organization= School of Computer Science,The University of Sydney , city= Sydney , postcode= NSW 2006 , state= NSW , country= Australia organization= School of Artificial Intelligence ,Anhui University , city= Hefei , postcode= 230601 , state= Anhui , country= China

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments This is a submitted version of a paper accepted by Information Fusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05637 2025-08-11 cs.HC cs.AI 57%

Automated Visualization Makeovers with LLMs

Siddharth Gangwar, David A. Selby, Sebastian J. Vollmer

机构 * University of Kaiserslautern–Landau (RPTU)(凯撒斯劳滕-兰道大学(RPTU)) Department of Data Science and its Applications, German Research Center for Artificial Intelligence (DFKI)(数据科学及其应用系,德国人工智能研究中心(DFKI))

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04389 2025-08-07 cs.AI 57%

GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement Learning

Weitai Kang, Bin Lei, Gaowen Liu, Caiwen Ding, Yan Yan

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) University of Minnesota(明尼苏达大学) Cisco Research(思科研究)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04239 2025-08-07 cs.CL 57%

DP-GPT4MTS: Dual-Prompt Large Language Model for Textual-Numerical Time Series Forecasting

Chanjuan Liu, Shengzhi Wang, Enqiang Zhu

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03998 2025-08-07 cs.CL 57%

Transferring Expert Cognitive Models to Social Robots via Agentic Concept Bottleneck Models

Xinyu Zhao, Zhen Tan, Maya Enisman, Minjae Seo, Marta R. Durantini, Dolores Albarracin, Tianlong Chen

机构 * The University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) Arizona State University(亚利桑那州立大学) University of Pennsylvania(宾夕法尼亚大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

Comments 27 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03345 2025-08-06 cs.AI 57%

Adaptive AI Agent Placement and Migration in Edge Intelligence Systems

Xingdan Wang, Jiayi He, Zhiqing Tang, Jianxiong Guo, Jiong Lou, Liping Qian, Tian Wang, Weijia Jia

机构 * Institute of Artificial Intelligence and Future Networks, Beijing Normal University, China(人工智能与未来网络研究院,北京师范大学,中国) Faculty of Arts and Sciences, Beijing Normal University, China(文理学院,北京师范大学,中国) Department of Computer Science and Engineering, Shanghai Jiao Tong University, China(计算机科学与工程系,上海交通大学,中国) College of Information Engineering, Zhejiang University of Technology, China(信息工程学院,浙江工业大学,中国)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01917 2025-08-05 cs.RO cs.AI 57%

L3M+P: Lifelong Planning with Large Language Models

Krish Agarwal, Yuqian Jiang, Jiaheng Hu, Bo Liu, Peter Stone

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01186 2025-08-05 cs.AI cs.HC 57%

A Survey on Agent Workflow -- Status and Future

Chaojia Yu, Zihan Cheng, Hanwen Cui, Yishuo Gao, Zexu Luo, Yijin Wang, Hangbin Zheng, Yong Zhao

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments 12 pages, 3 figures, accepted to IEEE Conference, ICAIBD(International Conference of Artificial Intelligence and Big Data) 2025. This is the author's version, not the publisher's. See https://ieeexplore.ieee.org/document/11082076

Journal ref IEEE ICAIBD 2025, pp. 770-781

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12974 2025-08-05 cs.CV cs.RO 57%

Exploring 3D Reasoning-Driven Planning: From Implicit Human Intentions to Route-Aware Activity Planning

Xueying Jiang, Wenhao Li, Xiaoqin Zhang, Ling Shao, Shijian Lu

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏