arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2773 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2773 篇

2512.19933 2025-12-24 cs.CL 70%

PRISM: A Personality-Driven Multi-Agent Framework for Social Media Simulation

PRISM: 一种基于个性的多智能体框架用于社交媒体模拟

Zhixiang Lu, Xueyuan Deng, Yiran Liu, Yulong Li, Qiang Yan, Imran Razzak, Jionglong Su

机构 * University of Liverpool(利物浦大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) University College London(伦敦大学学院) Xi'an Jiaotong-Liverpool University(西安交通大学-利物浦大学) Chinese Academy of Sciences(中国科学院) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CL

AI总结 PRISM通过结合连续情绪演变与基于个性的决策过程,提供了一种更准确模拟社交媒体中个性驱动意见极化的框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19280 2025-12-24 cs.CV 70%

RemoteReasoner: Towards Unifying Geospatial Reasoning Workflow

RemoteReasoner: 向统一地理空间推理流程迈进

Liang Yao, Fan Liu, Hongbo Lu, Chuanyi Zhang, Rui Min, Shengxiang Xu, Shimin Di, Pai Peng

机构 * COWARobot

专题命中 多模态Agent :multi-modal(abstract);MLLM(abstract);分类 cs.CV

AI总结 RemoteReasoner通过强化学习实现统一地理空间推理流程,具备多粒度任务处理能力和自主推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19865 2025-11-26 cs.AI 70%

Agentic AI-Empowered Conversational Embodied Intelligence Networks in 6G

基于代理AI的6G时代对话具身智能网络

Mingkai Chen, Zijie Feng, Lei Wang, Yaser Khamayseh

机构 * Key Laboratory of Broadband Wireless Communication and Sensor Network Technology(宽带无线通信与传感网络技术重点实验室) Nanjing University of Posts and Telecommunications(南京邮电大学) College of Technological Innovation(技术创新学院)

专题命中 多模态Agent :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出一种基于代理AI的6G时代对话具身智能网络,通过多模态融合、自适应通信和可解释性模块,提升复杂任务执行效率与语义一致性。

Comments 7 pages, 8 figures. Preprint submitted to IEEE Vehicle Technology Magazine

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06417 2025-11-11 cs.AI 70%

AUTO-Explorer: Automated Data Collection for GUI Agent

Xiangwu Guo, Difei Gao, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(新加坡国立大学展示实验室)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20462 2025-11-06 cs.AI 70%

TAMO: Fine-Grained Root Cause Analysis via Tool-Assisted LLM Agent with Multi-Modality Observation Data in Cloud-Native Systems

Xiao Zhang, Qi Wang, Mingyi Li, Yuan Yuan, Mengbai Xiao, Fuzhen Zhuang, Dongxiao Yu

机构 * School of Computer Science and Technology, Shandong University(山东大学计算机科学与技术学院) Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院)

专题命中 多模态Agent :multi-modal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27363 2025-11-03 cs.AI 70%

ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Use

Mengjie Deng, Guanting Dong, Zhicheng Dou

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21809 2025-10-28 cs.CV cs.RO 70%

Embodied Navigation with Auxiliary Task of Action Description Prediction

Haru Kondoh, Asako Kanezaki

机构 * Institute of Science Tokyo(东京科学研究所) RIKEN AIP(日本科学技术研究所AIP)

专题命中 多模态Agent :multimodal(abstract);audio-visual(abstract);分类 cs.CV

Comments ICCV 2025 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20838 2025-10-27 cs.AI cs.MA 70%

Sketch2BIM: A Multi-Agent Human-AI Collaborative Pipeline to Convert Hand-Drawn Floor Plans to 3D BIM

Abir Khan Ratul, Sanjay Acharjee, Somin Park, Md Nazmus Sakib

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09507 2025-10-13 cs.CV cs.RO 70%

PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs

Zixin Zhang, Kanghao Chen, Xingwang Lin, Lutao Jiang, Xu Zheng, Yuanhuiyi Lyu, Litao Guo, Yinchuan Li, Ying-Cong Chen

机构 * HKUST(GZ)(香港科技大学(广州)) HKUST(香港科技大学) Beihang University(北航大学) Knowin

专题命中 多模态Agent :multimodal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25185 2025-09-30 cs.CV 70%

PixelCraft: A Multi-Agent System for High-Fidelity Visual Reasoning on Structured Images

Shuoshuo Zhang, Zijian Li, Yizhen Zhang, Jingjing Fu, Lei Song, Jiang Bian, Jun Zhang, Yujiu Yang, Rui Wang

机构 * Microsoft Research(微软研究院) Tsinghua University(清华大学) Hong Kong University of Science and Technology(香港理工大学)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20499 2025-09-26 cs.RO cs.AI 70%

Boosting Zero-Shot VLN via Abstract Obstacle Map-Based Waypoint Prediction with TopoGraph-and-VisitInfo-Aware Prompting

Boqi Li, Siyuan Li, Weiyi Wang, Anran Li, Zhong Cao, Henry X. Liu

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12026 2025-09-26 cs.CV 70%

3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering

Rongtao Xu, Han Gao, Mingming Yu, Dong An, Shunpeng Chen, Changwei Wang, Li Guo, Xiaodan Liang, Shibiao Xu

专题命中 多模态Agent :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15635 2025-09-22 cs.AI 70%

MicroRCA-Agent: Microservice Root Cause Analysis Method Based on Large Language Model Agents

Pan Tang, Shixiang Tang, Huanqi Pu, Zhiqing Miao, Zhixing Wang

机构 * School of Communication and Information Engineering, Shanghai University, Shanghai, China(上海大学通信与信息工程学院) School of Communication and Electronic Engineering, East China Normal University, Shanghai, China(华东师范大学通信与电子工程学院) School of Information and Electronics, Beijing Institute of Technology, Beijing, China(北京理工大学信息与电子学院)

专题命中 多模态Agent :multimodal(abstract);cross-modal(abstract);分类 cs.AI

Comments 18 pages, 22 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11824 2025-09-18 cs.HC cs.AI 70%

AppAgent v2: Advanced Agent for Flexible Mobile Interactions

Yanda Li, Chi Zhang, Wenjia Jiang, Wanqi Yang, Bin Fu, Pei Cheng, Xin Chen, Ling Chen, Yunchao Wei

机构 * University of Technology Sydney(悉尼技术大学) Tencent(腾讯) Beijing Jiaotong University(北京交通大学) Westlake University(西湖大学)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01275 2025-09-04 cs.CV 70%

Novel Category Discovery with X-Agent Attention for Open-Vocabulary Semantic Segmentation

Jiahao Li, Yang Lu, Yachao Zhang, Fangyong Wang, Yuan Xie, Yanyun Qu

专题命中 多模态Agent :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by ACMMM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13073 2025-09-04 cs.RO cs.CV 70%

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey

Rui Shao, Wei Li, Lingsen Zhang, Renshan Zhang, Zhiyang Liu, Ran Chen, Liqiang Nie

机构 * School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen)(计算机科学与技术学院,哈尔滨工业大学(深圳))

专题命中 多模态Agent :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments Project Page: https://github.com/JiuTian-VL/Large-VLM-based-VLA-for-Robotic-Manipulation

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10414 2025-08-13 eess.IV cs.CV 70%

Style transfer between Microscopy and Magnetic Resonance Imaging via Generative Adversarial Network in small sample size settings

Monika Pytlarz, Adrian Onicas, Alessandro Crimi

机构 * Sano – Centre for Computational Personalised Medicine(Sano 个性化医学计算中心)

专题命中 多模态Agent :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments 2023 IEEE International Conference on Image Processing (ICIP)

Journal ref 2023 IEEE International Conference on Image Processing (ICIP), Kuala Lumpur, Malaysia, 2023, pp. 1120-1124

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17462 2025-07-24 cs.CV 70%

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents

Chang Nie, Guangming Wang, Zhe Lie, Hesheng Wang

机构 * School of Automation and Intelligent Sensing, Shanghai Jiao Tong University and Key Laboratory of System Control and Information Processing, Ministry of Education of China(自动化与智能感知学院,上海交通大学;系统控制与信息处理重点实验室,中华人民共和国教育部)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01376 2025-07-03 cs.AI 70%

AI Agents and Agentic AI-Navigating a Plethora of Concepts for Future Manufacturing

Yinwang Ren, Yangyang Liu, Tang Ji, Xun Xu

机构 * Department of Mechanical and Mechatronics Engineering, Faculty of Engineering and Design, University of Auckland(机械与机电工程系,工程与设计学院,奥克兰大学)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.AI

Comments Submitted to JMS(March 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17612 2025-06-24 cs.CV 70%

JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent

Yunlong Lin, Zixu Lin, Kunjie Lin, Jinbin Bai, Panwang Pan, Chenxin Li, Haoyu Chen, Zhongdao Wang, Xinghao Ding, Wenbo Li, Shuicheng Yan

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(中国教育部多媒体可信感知与高效计算重点实验室,厦门大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Chinese University of Hong Kong(香港中文大学) Bytedance(字节跳动) National University of Singapore(新加坡国立大学) Tsinghua University(清华大学)

专题命中 多模态Agent :multi-modal(abstract);MLLM(abstract);分类 cs.CV

Comments 40 pages, 26 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10069 2025-06-18 cs.RO cs.CV 70%

SmartWay: Enhanced Waypoint Prediction and Backtracking for Zero-Shot Vision-and-Language Navigation

Xiangyu Shi, Zerui Li, Wenqi Lyu, Jiatong Xia, Feras Dayoub, Yanyuan Qiao, Qi Wu

机构 * Australian Institute for Machine Learning at the University of Adelaide(澳大利亚机器学习研究所(阿德莱德大学))

专题命中 多模态Agent :multi-modal(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by IROS 2025. Project website: https://sxyxs.github.io/smartway/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17399 2025-05-27 cs.CL 70%

FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow

Haoyu Sun, Huichen Will Wang, Jiawei Gu, Linjie Li, Yu Cheng

机构 * Tongji University(同济大学) University of Washington(华盛顿大学) Sun Yat-sen University(中山大学) Microsoft(微软) The Chinese University of Hong Kong(香港中文大学)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21620 2025-05-27 cs.AI 70%

UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning

Zhengxi Lu, Yuxiang Chai, Yaxuan Guo, Xi Yin, Liang Liu, Hao Wang, Han Xiao, Shuai Ren, Guanjing Xiong, Hongsheng Li

机构 * vivo AI Lab(vivo人工智能实验室)

专题命中 多模态Agent :multimodal(abstract);multi-modal(abstract);分类 cs.AI

Comments Updated UI-R1-E-3B

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20464 2025-05-14 cs.AI 70%

A Survey on GUI Agents with Foundation Models Enhanced by Reinforcement Learning

Jiahao Li, Kaer Huang

机构 * Lenovo Research(联想研究院)

专题命中 多模态Agent :multimodal(abstract);multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01268 2024-12-03 cs.CV 70%

Ponder & Press: Advancing Visual GUI Agent towards General Computer Control

Yiqin Wang, Haoji Zhang, Jingqi Tian, Yansong Tang

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10252 2024-11-18 cs.CV 70%

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning

Jingru Yang, Huan Yu, Yang Jingxin, Chentianye Xu, Yin Biao, Yu Sun, Shengfeng He

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16048 2024-10-29 cs.HC cs.AI 70%

GUIDE: Graphical User Interface Data for Execution

Rajat Chawla, Adarsh Jha, Muskaan Kumar, Mukunda NS, Ishaan Bhola

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.AI

Comments 11 pages, 8 figures, 3 Tables and 1 Algorithm

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11872 2024-10-18 cs.HC cs.AI cs.LG 70%

ClickAgent: Enhancing UI Location Capabilities of Autonomous Agents

Jakub Hoscilowicz, Bartosz Maj, Bartosz Kozakiewicz, Oleksii Tymoshchuk, Artur Janicki

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.AI

Comments The code for ClickAgent is available at github.com/Samsung/ClickAgent

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14170 2024-09-24 cs.CV 70%

LFP: Efficient and Accurate End-to-End Lane-Level Planning via Camera-LiDAR Fusion

Guoliang You, Xiaomeng Chu, Yifan Duan, Xingchen Li, Sha Zhang, Jianmin Ji, Yanyong Zhang

专题命中 多模态Agent :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17842 2024-07-26 cs.LG cs.AI 70%

On the Opportunities of (Re)-Exploring Atmospheric Science by Foundation Models: A Case Study

Lujia Zhang, Hanzhe Cui, Yurong Song, Chenyue Li, Binhang Yuan, Mengqian Lu

专题命中 多模态Agent :multimodal(abstract);multimodal foundation model(abstract);分类 cs.AI

Comments 28 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏