arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2766 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2766 篇

2602.07215 2026-02-10 eess.SY cs.AI cs.DC cs.SY 79%

Multi-Agentic AI for Fairness-Aware and Accelerated Multi-modal Large Model Inference in Real-world Mobile Edge Networks

多智能体AI用于现实世界移动边缘网络中公平性感知和加速的多模态大模型推理

Haiyuan Li, Hari Madhukumar, Shuangyi Yan, Yulei Wu, Dimitra Simeonidou

机构 * Smart Internet Lab, Department of Electrical and Electronic Engineering, University of Bristol(布里斯托大学电子与电气工程系智能互联网实验室)

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.AI

AI总结 本文提出多智能体AI框架,通过优化提示路由和模型部署,实现移动边缘网络中多模态大模型推理的低延迟与高公平性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04763 2026-02-05 cs.LG cs.AI 79%

Active Asymmetric Multi-Agent Multimodal Learning under Uncertainty

在不确定性下主动非对称多智能体多模态学习

Rui Liu, Pratap Tokekar, Ming Lin

机构 * University of Maryland, College Park(马里兰大学 College Park分校)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 在不确定性下主动非对称多智能体多模态学习通过模态层面协作提升事故检测性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04112 2026-02-05 cs.CL 79%

DELTA: Deliberative Multi-Agent Reasoning with Reinforcement Learning for Multimodal Psychological Counseling

DELTA: 基于强化学习的多智能体推理用于多模态心理辅导

Jiangnan Yang, Junjie Chen, Fei Wang, Yiqi Nie, Yuxin Liu, Zhangling Duan, Jie Chen

机构 * Anhui University(安徽大学) Hefei University of Technology(合肥工业大学) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合国家科学中心人工智能研究所)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL

AI总结 DELTA 通过强化学习和多智能体推理提升多模态心理辅导的质量与情感契合度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00993 2026-02-03 cs.RO cs.AI 79%

HERMES: A Holistic End-to-End Risk-Aware Multimodal Embodied System with Vision-Language Models for Long-Tail Autonomous Driving

HERMES: 一种集成端到端风险感知多模态具身系统,用于长尾自动驾驶

Weizhe Tang, Junwei You, Jiaxi Liu, Zhaoyi Wang, Rui Gan, Zilin Huang, Feng Wei, Bin Ran

机构 * Department of Civil and Environmental Engineering, University of Wisconsin–Madison(土木与环境工程系,威斯康星大学麦迪逊分校)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 HERMES通过整合视觉-语言模型和多模态感知,提升自动驾驶在长尾混合交通场景中的风险感知和轨迹规划能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00454 2026-02-03 cs.AI 79%

Cross-Modal Memory Compression for Efficient Multi-Agent Debate

多模态记忆压缩以实现高效的多智能体辩论

Jing Wu, Yue Sun, Tianpei Xie, Suiyao Chen, Jingyuan Bao, Yaopengxiao Xu, Gaoyuan Du, Inseok Heo, Alexander Gutfraind, Xin Wang

专题命中 多模态Agent :cross-modal(title,abstract);分类 cs.AI

AI总结 DebateOCR通过多模态压缩框架将多智能体辩论的历史文本转换为图像表示,显著减少token使用并提升推理效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21403 2026-01-30 cs.AI cs.MA 79%

DataCross: A Unified Benchmark and Agent Framework for Cross-Modal Heterogeneous Data Analysis

DataCross: 一个统一的基准和代理框架,用于跨模态异构数据分析

Ruyi Qi, Zhou Liu, Wentao Zhang

机构 * Peking University(北京大学)

专题命中 多模态Agent :cross-modal(title,abstract);分类 cs.AI

AI总结 DataCross提出一个统一的基准和代理框架,用于跨模态异构数据分析,通过多步骤联合推理提升事实性与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20676 2026-01-29 cs.CL 79%

Efficient Multimodal Planning Agent for Visual Question-Answering

高效的多模态规划代理用于视觉问答

Zhuo Chen, Xinyu Geng, Xinyu Wang, Yong Jiang, Zhen Zhang, Pengjun Xie, Kewei Tu

机构 * ShanghaiTech University(上海科技大学) Alibaba Tongyi Lab(阿里云通义实验室)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL

AI总结 本文提出了一种高效的多模态规划代理,通过动态分解mRAG流程,提升视觉问答任务的效率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03709 2026-01-28 cs.HC cs.AI 79%

AutoGameUI: Constructing High-Fidelity GameUI via Multimodal Correspondence Matching

AutoGameUI: 通过多模态对应匹配构建高保真的游戏界面

Zhongliang Tang, Qingrong Cheng, Mengchen Tan, Yongxiang Zhang, Fei Xia

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 AutoGameUI通过多模态对应匹配自动构建高保真的游戏界面,提升了游戏界面设计的效率和一致性。

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10046 2026-01-27 cs.AI 79%

SimWorld-Robotics: Synthesizing Photorealistic and Dynamic Urban Environments for Multimodal Robot Navigation and Collaboration

SimWorld-Robotics: 为多模态机器人导航与协作合成逼真动态城市环境

Yan Zhuang, Jiawei Ren, Xiaokang Ye, Jianzhi Shen, Ruixuan Zhang, Tianai Yue, Muhammad Faayez, Xuhong He, Ziqiao Ma, Lianhui Qin, Zhiting Hu, Tianmin Shu

机构 * University of Virginia(弗吉尼亚大学) UC San Diego(加州大学圣地亚哥分校) Johns Hopkins University(约翰霍普金斯大学) Carnegie Mellon University(卡内基梅隆大学) University of Michigan(密歇根大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 SimWorld-Robotics通过合成逼真动态城市环境,提出两个多模态机器人基准测试,评估机器人在复杂场景中的导航、协作与通信能力。

Comments Conference: NeurIPS 2025 (main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13976 2026-01-26 cs.CV cs.RO 79%

FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-Language Navigation

FantasyVLN: 一种统一的多模态链式推理框架用于视觉-语言导航

Jing Zuo, Lingzhou Mu, Fan Jiang, Chengcheng Ma, Mu Xu, Yonggang Qi

机构 * Fantasy AIGC Team(幻想AIGC团队) Beijing University of Posts and Telecommunications(北京邮电大学) Tsinghua University(清华大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

AI总结 FantasyVLN通过统一的隐式推理框架实现视觉-语言导航的实时高效推理,结合多模态信息提升导航性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01910 2026-01-23 cs.AI 79%

MMP-A*: Multimodal Perception Enhanced Incremental Heuristic Search on Path Planning

MMP-A*:多模态感知增强的路径规划增量启发式搜索

Minh Hieu Ha, Khanh Ly Ta, Hung Phan, Tung Doan, Tung Dao, Dao Tran, Huynh Thi Thanh Binh

机构 * Hanoi University of Science and Technology(河内科学技术大学) Vingroup Big Data Research Center(VinGroup大数据研究中心) FPT Software AI Center(FPT软件AI中心)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 MMP-A*通过融合视觉语言模型的空间感知与自适应衰减机制,提升路径规划的几何精度与计算效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12981 2026-01-21 cs.CV cs.LG 79%

Early Prediction of Type 2 Diabetes Using Multimodal data and Tabular Transformers

利用多模态数据和表格变压器进行2型糖尿病早期预测

Sulaiman Khan, Md. Rafiul Biswas, Zubair Shah

机构 * College of Science and Engineering(科学与工程学院)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

AI总结 利用TabTrans分析多模态数据,通过预测T2DM风险,提升糖尿病管理的前瞻性与个性化水平。

Comments 08 pages, 06 figures, accepted for publication in FLLM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12037 2026-01-21 cs.HC cs.CV cs.SY eess.SY 79%

Multimodal Feedback for Handheld Tool Guidance: Combining Wrist-Based Haptics with Augmented Reality

多模态反馈用于手持工具引导:结合基于手腕的触觉反馈与增强现实

Yue Yang, Christoph Leuze, Brian Hargreaves, Bruce Daniel, Fred M Baik

机构 * Stanford University(斯坦福大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

AI总结 本研究通过结合增强现实与触觉反馈,提升手术中手持工具的精确引导与操作效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08434 2026-01-21 cs.RO cs.AI 79%

Large Multimodal Models for Embodied Intelligent Driving: The Next Frontier in Self-Driving?

大规模多模态模型用于具身智能驾驶:自我驾驶的下一个前沿?

Long Zhang, Yuchen Xia, Bingqing Wei, Zhen Liu, Shiwen Mao, Zhu Han, Mohsen Guizani

机构 * School of Information and Electrical Engineering, Hebei University of Engineering(河北工程大学信息与电子工程学院) School of Information Science and Engineering, Lanzhou University(兰州大学信息科学与工程学院) Department of Electrical and Computer Engineering, Auburn University(阿肯色大学电气与计算机工程系) Department of Electrical and Computer Engineering, University of Houston(休斯顿大学电气与计算机工程系) Department of Computer Science and Engineering, Kyung Hee University(庆熙大学计算机科学与工程系) Machine Learning Department, Mohamed Bin Zayed University of Artificial Intelligence(Mohamed Bin Zayed人工智能大学机器学习系)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出了一种语义和策略双驱动的混合决策框架,用于提升具身智能驾驶中的持续学习与联合决策能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18177 2026-01-21 cs.CV cs.LG cs.MA 79%

Scene-Aware Vectorized Memory Multi-Agent Framework with Cross-Modal Differentiated Quantization VLMs for Visually Impaired Assistance

具有跨模态差异化量化功能的场景感知向量内存多智能体框架用于视障辅助

Xiangxiang Wang, Xuanyu Wang, YiJia Luo, Yongbin Yu, Manping Fan, Jingtao Zhang, Liyong Ren

机构 * School of Information Software Engineering, University of Electronic Science Sichuan Provincial Key Laboratory for Human Disease Gene Study, Sichuan Academy of Medical Sciences \& Sichuan Provincial People's Hospital, University of Electronic Science Faculty of Computing, Harbin Institute of Technology, Harbin, China

专题命中 多模态Agent :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出了一种具有跨模态差异化量化的多智能体框架,通过减少内存消耗和提升处理效率,为视障人士提供更高效的环境感知与辅助导航支持。

Comments 28 pages,9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24385 2026-01-09 cs.CV cs.RO 79%

Forging Spatial Intelligence: A Roadmap of Multi-Modal Data Pre-Training for Autonomous Systems

锻造空间智能:面向自主系统多模态数据预训练的路线图

Song Wang, Lingdong Kong, Xiaolu Liu, Hao Shi, Wentong Li, Jianke Zhu, Steven C. H. Hoi

机构 * Zhejiang University(浙江大学) National University of Singapore(新加坡国立大学) Nanjing University of Aeronautics and Astronautics(南京航空航天大学) Alibaba Group(阿里巴巴集团) Singapore Management University(新加坡管理学院)

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出多模态数据预训练框架,旨在通过整合多传感器数据提升自主系统空间智能,解决单模态模型整合难题,并提出统一的预训练范式分类及未来发展方向。

Comments Survey; 40 pages, 7 figures, 9 tables; GitHub Repo at https://github.com/worldbench/awesome-spatial-intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23412 2026-01-08 cs.AI 79%

MindWatcher: Toward Smarter Multimodal Tool-Integrated Reasoning

MindWatcher:迈向更智能的多模态工具集成推理

Jiawei Chen, Xintian Shen, Lihao Zheng, Zhenwei Shao, Handong Cui, Chaoqun Du, Li Gong, Feng Gu, Xuefeng Hao, Wei He, Jiabang He, Yi Hu, Bin Huang, Shanshan Li, Qizhen Li, Jing Luo, Zide Liu, Xiaobo Liu, Ning Mao, Lifu Mu, Xuhao Pan, Zhiheng Qu, Chang Ren, Xudong Rao, Haoyi Sun, Qian Wang, Shuai Wang, Zhichao Wang, Wei Wang, Lian Wen, Jiqing Zhan, Hongfu Yang, Sheng Yang, Jiajun Yang, Pengfei Yu, Hongyuan Zhang, Bin Zhang, Chunpeng Zhou, Zheng Zhou, Shucheng Zhou, Shuo Xie, Yun Zhu, Hao Ma, Tao Wei, Pan Zhou, Wei Chen

机构 * Li Auto Inc(力汽车公司)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 MindWatcher是一种能够自主调用工具并进行多模态推理的智能体,通过高效训练和高质量数据集提升了多步骤决策任务的性能。

Comments Technique Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01366 2026-01-06 cs.AI 79%

KGCE: Knowledge-Augmented Dual-Graph Evaluator for Cross-Platform Educational Agent Benchmarking with Multimodal Language Models

KGCE:基于多模态语言模型的跨平台教育代理基准评估知识增强双图评估器

Zixian Liu, Sihao Liu, Yuqi Zhao

机构 * Faculty of the School of Computer Science, Central China Normal University(中央财经大学计算机学院)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 KGCE提出一种基于多模态语言模型的跨平台教育代理基准评估框架,通过知识库增强和双图评估方法提升对特定学校软件任务的执行效率和评估精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00730 2026-01-05 cs.CV 79%

Grading Handwritten Engineering Exams with Multimodal Large Language Models

用多模态大语言模型评分手写工程考试

Janez Perš, Jon Muhovič, Andrej Košir, Boštjan Murovec

机构 * University of Ljubljana, Faculty of Electrical Engineering(卢布尔雅那大学电子工程学院)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

AI总结 本研究利用多模态大语言模型对斯洛文尼亚语手写工程考试进行自动评分,通过多阶段设计实现高可靠性,达到与人工评分约8分的均方差,验证了结构化提示和参考定位的重要性。

Comments 10 pages, 5 figures, 2 tables. Supplementary material available at https://lmi.fe.uni-lj.si/en/janez-pers-2/supplementary-material/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22998 2026-01-01 cs.AI 79%

TIM-PRM: Verifying multimodal reasoning with Tool-Integrated PRM

TIM-PRM:利用工具集成PRM验证多模态推理

Peng Kuang, Xiangxiang Wang, Wentao Liu, Jian Dong, Kaidi Xu

机构 * Zhejiang University(浙江大学) iFLYTEK AI Research Institute(iFLYTEK人工智能研究院) City University of Hong Kong(香港城市大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 TIM-PRM通过工具集成PRM验证多模态推理,有效解决视觉幻觉和逻辑不一致问题,实验表明其在性能和可解释性上优于现有模型。

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13256 2025-12-24 cs.AI cs.CY cs.MA 79%

CardAIc-Agents: A Multimodal Framework with Hierarchical Adaptation for Cardiac Care Support

CardAIc-Agents: 一种用于心脏护理支持的多模态框架,具有分层适应性

Yuting Zhang, Karina V. Bunting, Asgher Champsi, Xiaoxia Wang, Wenqi Lu, Alexander Thorley, Sandeep S Hothi, Zhaowen Qiu, Baturalp Buyukates, Dipak Kotecha, Jinming Duan

机构 * School of Computer Science, University of Birmingham, UK(英国伯明翰大学计算机科学学院) Department of Cardiovascular Sciences, University of Birmingham, UK(英国伯明翰大学心血管科学系) NIHR Birmingham Biomedical Research Centre and West Midlands NHS Secure Data Environment, University Hospitals Birmingham NHS Foundation Trust, UK(英国伯明翰大学医院 NHS 基础信任机构) Department of Computing and Mathematics, Manchester Metropolitan University, UK(曼彻斯特 Metropolitan 大学计算与数学系) Department of Cardiology, Heart and Lung Centre, Royal Wolverhampton NHS Trust, UK(皇家沃尔夫汉普顿 NHS 委员会心内科部门) College of Computer and Control Engineering, Northeast Forestry University, China(中国东北林业大学计算机与控制工程学院) Julius Center, University Medical Center Utrecht, the Netherlands(荷兰乌得勒支大学医学中心朱利叶斯中心) Division of Informatics, Imaging and Data Sciences, University of Manchester, UK(英国曼彻斯特大学信息学、成像与数据科学系)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 CardAIc-Agents通过多模态框架和分层适应性,提升心脏护理支持的效率和灵活性,有效解决传统AI代理在适应性推理、工具支持、知识更新和视觉输出方面的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19000 2025-12-23 cs.HC cs.AI cs.SY eess.SY 79%

An AI-Driven Multimodal Smart Home Platform for Continuous Monitoring and Assistance in Post-Stroke Motor Impairment

基于人工智能的多模态智能家庭平台:用于中风后运动功能障碍的持续监测与协助

Chenyu Tang, Ruizhi Zhang, Shuo Gao, Zihe Zhao, Zibo Zhang, Jiaqi Wang, Cong Li, Junliang Chen, Yanning Dai, Shengbo Wang, Ruoyu Juan, Qiaoying Li, Ruimou Xie, Xuhang Chen, Xinkai Zhou, Yunjia Xia, Jianan Chen, Fanghao Lu, Xin Li, Ninglli Wang, Peter Smielewski, Yu Pan, Hubin Zhao, Luigi G. Occhipinti

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出一种基于人工智能的多模态智能家庭平台,通过整合可穿戴设备和环境传感器,实现中风后患者运动功能的持续监测与智能协助,显著提升用户满意度。

Comments 5 figures, 41 references

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20164 2025-12-16 cs.CL 79%

Thinking with Visual Abstract: Enhancing Multimodal Reasoning via Visual Abstraction

通过视觉抽象思考:通过视觉抽象增强多模态推理

Dairu Liu, Ziyue Wang, Minyuan Ruan, Fuwen Luo, Chi Chen, Peng Li, Yang Liu

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL

AI总结 通过引入视觉抽象思考范式,提升多模态大语言模型在视觉感知和推理任务中的性能,实现更高效的视觉推理机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10975 2025-12-15 cs.LG cs.AI cs.HC cs.MA 79%

Agent-Based Modular Learning for Multimodal Emotion Recognition in Human-Agent Systems

基于代理的模块化学习:人类-代理系统中多模态情感识别

Matvey Nepomnyaschiy, Oleg Pereziabov, Anvar Tliamov, Stanislav Mikhailov, Ilya Afanasyev

机构 * Research Center of the Artificial Intelligence Institute, Innopolis University(人工智能研究所研究中心,因诺波利斯大学) ITMO University(ITMO大学) Saint Petersburg Electrotechnical University ``LETI''(圣彼得堡电工大学『LETI』)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出了一种基于代理的模块化学习框架,用于多模态情感识别,通过自主代理协调实现灵活、高效的训练和维护。

Comments 14 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02530 2025-12-10 cs.AI 79%

Aetheria: A multimodal interpretable content safety framework based on multi-agent debate and collaboration

Aetheria:基于多智能体辩论与协作的多模态可解释内容安全框架

Yuxiang He, Jian Zhao, Yuchen Yuan, Tianle Zhang, Wei Cai, Haojie Cheng, Ziyan Shi, Ming Zhu, Haichuan Tang, Chi Zhang, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信) Sichuan University(四川大学) Peking University(北京大学) Beijing Jiaotong University(北京交通大学) Harbin Institute of Technology(哈尔滨工业大学) China Railway Rolling Stock Corporation Limited(中国铁路滚动股票有限公司)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 Aetheria通过多智能体辩论与协作机制,实现多模态内容安全的可解释性与高准确性,提升可信AI审查水平。

Comments https://github.com/Herrieson/Aetheria

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02395 2025-12-09 cs.CV 79%

Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch

Skywork-R1V4:通过图像与深度研究交织思考实现代理多模态智能

Yifan Zhang, Liang Hu, Haofeng Sun, Peiyu Wang, Yichen Wei, Shukang Yin, Jiangbo Pei, Wei Shen, Peng Xia, Yi Peng, Tianyidan Xie, Eric Li, Yang Liu, Xuchen Song, Yahui Zhou

机构 * Skywork AI

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

AI总结 Skywork-R1V4通过交织推理实现多模态代理智能,仅用监督学习在少数据上训练,超越现有模型在多个基准测试中的表现。

Comments 21 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05111 2025-12-05 cs.CV 79%

ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning

ARM-Thinker: 通过智能工具使用和视觉推理强化多模态生成奖励模型

Shengyuan Ding, Xinyu Fang, Ziyu Liu, Yuhang Zang, Yuhang Cao, Xiangyu Zhao, Haodong Duan, Xiaoyi Dong, Jianze Liang, Bin Wang, Conghui He, Dahua Lin, Jiaqi Wang

机构 * Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Zhejiang University(浙江大学) Shanghai Jiao Tong University(上海交通大学) The Chinese University of Hong Kong(香港中文大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

AI总结 ARM-Thinker通过智能工具使用和视觉推理提升多模态奖励模型的准确性与可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12679 2025-12-05 cs.CV 79%

TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents

TongUI: 通过多模态网络教程构建大规模GUI代理

Bofei Zhang, Zirui Shang, Zhi Gao, Wang Zhang, Rui Xie, Xiaojian Ma, Tao Yuan, Xinxiao Wu, Song-Chun Zhu, Qing Li

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

AI总结 TongUI通过多模态网络教程构建大规模GUI代理,利用GUI-Net数据集提升定位和导航性能,优于基线代理10%。

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20085 2025-12-04 cs.AI cs.MA 79%

VICoT-Agent: A Vision-Interleaved Chain-of-Thought Framework for Interpretable Multimodal Reasoning and Scalable Remote Sensing Analysis

VICoT-Agent:一种用于可解释多模态推理和可扩展遥感分析的视觉交织思维链框架

Chujie Wang, Zhiyuan Luo, Ruiqi Liu, Can Ran, Shenghua Fan, Xi Chen, Chu He

机构 * Wuhan University(武汉大学) University of Toronto(多伦多大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 VICoT-Agent通过视觉交织思维链框架实现多模态推理和遥感分析,采用堆栈结构和模块化工具集提升推理效率,并通过推理堆栈蒸馏方法降低模型复杂度,显著提升推理透明度和执行效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17425 2025-11-21 cs.AI 79%

Evaluating Multimodal Large Language Models with Daily Composite Tasks in Home Environments

在家庭环境中评估多模态大语言模型的日常复合任务

Zhenliang Zhang, Yuxi Wang, Hongzhao Xie, Shiyun Zhao, Mingyuan Liu, Yujie Lu, Xinyi He, Zhenku Cheng, Yujia Peng

机构 * State Key Laboratory of General Artificial Intelligence, Beijing Institute for General Artificial Intelligence(通用人工智能国家重点实验室、北京通用人工智能研究院) School of Psychological and Cognitive Sciences and Beijing Key Laboratory of Behavior and Mental Health, Key Laboratory of Machine Perception (Ministry of Education), Peking University(心理与认知科学学院及北京行为与心理健康重点实验室、机器感知重点实验室(教育部)) School of Intelligence Science and Technology, Peking University(智能科学与技术学院)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 本研究在家庭环境中评估了多模态大语言模型在日常复合任务中的表现,发现其在物体理解、空间智能和社会活动领域存在显著差距,为具身MLLMs的发展提供了初步评估框架。

详情

展开后加载摘要…

URL PDF HTML 收藏