arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-08-21 至 2026-08-21 共收录 11 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 11 篇

2608.19739 2026-08-21 cs.CV cs.AI cs.LG 新提交 84%

Question-Guided Evidence Acquisition for Multimodal Visual Question Answering

面向多模态视觉问答的问题引导式证据获取

Alin-Ionut Popa

机构 * Amazon Inc.(亚马逊公司)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 针对多模态视觉问答中模型读取文档不可靠的问题,提出Q-Guide智能体引导感知,在两个数据集上优于基线方法,且增益来自精准感知引导而非复杂控制逻辑。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20237 2026-08-21 cs.AI 新提交 79%

Rule-Compliant Visual Spatial Planning for Multimodal Large Language Models

面向多模态大语言模型的符合规则的视觉空间规划

Yu Chen, Ting Lei, Yaoyi Li, Jia Cai, Zhecen Wu, Yang Liu

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王选计算机研究所) Yinwang Intelligent Technology Co., Ltd(银湾智能科技有限公司)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 该研究针对多模态大语言模型的规则遵循空间规划问题,构建了RuleMaze基准,提出语言-逻辑-函数混合方法和解耦多模态规划(DMP),提升了规则遵循度与规划成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.13854 2026-08-21 cs.CL 版本更新 79%

SPyCE: Skill-Policy Co-evolution for Multimodal Agents

SPyCE:多模态智能体的技能-策略协同进化

Ru Zhang, Weijie Qiu

机构 * Zhejiang University(浙江大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL

AI总结 研究多模态智能体,提出SPyCE框架,将推理轨迹提炼为分层技能库与策略协同进化,执行技能捕获局部操作,工作流技能编码高级先验,实验证明该框架优于基线,为构建多模态智能体提供新范式。

Comments Accepted by ACM MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19537 2026-08-21 cs.RO 新提交 78%

Multimodal Trajectory Planning for Surface Vehicles using Turning Circle-based Control Barrier Functions

基于转向圆型控制障碍函数的水面航行器多模态轨迹规划

Changyu Lee

机构 * Kongju National University(公州国立大学)

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 本文针对动态环境下自主水面航行器,提出结合MPC与TC-CBF的无引导路径多模态轨迹规划框架,可提升避碰成功率、减少安全违规并缓解局部极小问题。

Comments This work has been submitted to an Elsevier journal for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02794 2026-08-21 cs.AI 版本更新 77%

CharTool: Tool-Integrated Visual Reasoning for Chart Understanding

CharTool: 集成工具的视觉推理用于图表理解

Situo Zhang, Yifan Zhang, Zichen Zhu, Da Ma, Lei Pan, Danyang Zhang, Zihan Zhao, Lu Chen, Kai Yu

机构 * X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学与工程学院X-LANCE实验室) Jiangsu Key Lab of Language Computing(江苏省语言计算重点实验室) Suzhou Laboratory(苏州实验室) AISpeech Co., Ltd.(思必驰科技股份有限公司)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.AI

AI总结 本文提出CharTool,通过集成工具提升多模态大语言模型对图表的理解能力,通过双重数据管道和代理强化学习,在六个图表基准测试中取得显著提升。

Comments Accepted by ACMMM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20320 2026-08-21 cs.AI cs.CL 新提交 62%

An Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Prediction

一种用于主动数据收集、出行行为建模及天气敏感需求预测的智能体方法

Narges Ahmadi (1), Yubo Jiao (1), Jônatas Augusto Manzolli (1), Jiangbo Yu (1), Luis Miranda-Moreno (1) ((1) McGill University)

机构 * McGill University(麦吉尔大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本研究提出三智能体工作流,结合对话式调查、传统建模与多模态LLM预测,基于学生通勤数据实现天气敏感出行需求预测,视觉配置模型准确率达71.5%

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20129 2026-08-21 cs.MA cs.CL cs.CV 新提交 62%

Multi-Agent Orchestration with the Common-Sense Reasoning Capabilities of LLMs for Autonomous Driving

结合大语言模型常识推理能力的多智能体协同框架用于自动驾驶

Mehdi Azarafza, Faezeh Pasandideh, Ali Ehteshami Bejnordi, Stefan Henkler, Achim Rettberg

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 该研究针对自动驾驶中强化学习等方法的上下文推理缺陷,提出结合LLM常识推理的混合多智能体协同框架,经CARLA场景验证可保留结构化控制与安全机制,具备应用潜力。

Comments 17 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12616 2026-08-21 cs.AI cs.MM 版本更新 62%

Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs

每幅画都讲述一个危险的故事:基于内存增强的多智能体 jailbreak 攻击 VLMs

Jianhao Chen, Shiqin Wang, Haoyang Chen, Hanjie Zhao, Haozhe Liang, Zheng Wang, Tieyun Qian

机构 * Wuhan University(武汉大学) Zhongguancun Academy(中关村学院) Tianjin University(天津大学) University of the Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI、cs.MM

AI总结 本文提出MemJack框架,通过视觉语义引导多智能体协作,利用迭代空域投影过滤器提升对VLMs的攻击成功率,实验表明其在COCO数据集上达到71.48%的攻击成功率。

Comments This work was accepted by WISE2026. 15 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19794 2026-08-21 cs.AI cs.RO 新提交 57%

Towards general embodied intelligence: integrating large language models, knowledge bases, and reasoning capabilities to build the next generation of AI agents

迈向通用具身智能:整合大语言模型、知识库与推理能力以构建下一代AI智能体

Fujiang Yuan, Xia Huang, Lusheng Wang, Jun Ding, Zhen Tian, Yuxin Wang, Shaojie Gu, Yuki Funabora, Yanhong Peng, Zebing Mao

机构 * College of Mechanical Engineering, Chongqing University of Technology(重庆理工大学机械工程学院) James Watt School of Engineering, University of Glasgow(格拉斯哥大学詹姆斯·瓦特工程学院) School of Energy and Power, Jiangsu University of Science and Technology(江苏科技大学能源与动力学院) Magnesium Research Center, Kumamoto University(熊本大学镁研究中心) Department of Information and Communication Engineering, Nagoya University(名古屋大学信息与通信工程系) State Key Laboratory of Fluid Power and Mechatronic Systems, Zhejiang University(浙江大学流体动力与机电系统国家重点实验室)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 本文综述以LLM为核心的智能系统演进,提出整合LLMs、KBs、RA与具身性的概念框架,明确高效LLM部署等五大挑战,为开发复杂动态环境下的自适应多模态智能体提供路线图。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11665 2026-08-21 cs.CV cs.RO 版本更新 57%

UAV-Based Infrastructure Inspections: A Literature Review and Proposed Framework for AEC+FM

基于无人机的基础设施检查:面向AEC+FM的文献综述与提出框架

Amir Farzin Nikkhah, Dong Chen, Bradford Campbell, Somayeh Asadi, Arsalan Heydarian

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 本文综述了无人机在基础设施检查中的应用,提出框架整合多模态数据与Transformer架构,以提高检测准确性和可靠性,未来研究方向包括轻量级AI模型和合成数据集。

Comments Accepted for publication in the Proceedings of the International Conference on Computing in Civil Engineering (i3CE 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19964 2026-08-21 cs.LG 新提交 50%

G-MARK: Grounded Multi-Agent Reasoning for Cooperative Driving via Knowledge Graphs

G-MARK:基于知识图谱的协同驾驶接地多智能体推理

Bhavya Gupta, Onat Gungor, Tajana Rosing

机构 * University of California, San Diego(加利福尼亚大学圣迭戈分校) West Virginia University(西弗吉尼亚大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 提出 G-MARK 框架,通过知识图谱实现协同驾驶多智能体推理,提升遮挡推理与控制选择性能,减小通信负载,效果优于现有基线。

Comments Accepted for oral presentation at the 25th IEEE International Conference on Machine Learning and Applications (ICMLA'26)

详情

展开后加载摘要…

URL PDF HTML 收藏